S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
2025/05/30 by Xie, Yichen, Xu, Runsheng, He, Tong +9 · 8 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2505.24139
Abstract
The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches--which directly learn from sensor inputs to generate planning trajectories without human annotations often underperform the state of the art. We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. To this end, we propose S4-Driver, a scalable self-supervised motion planning algorithm with spatio-temporal visual representation, based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs.
Citations
- PaliGemma: A versatile 3B VLM for transfer
- Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
- When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
- MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- GenAD: Generative End-to-End Autonomous Driving
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
- DriveLM: Driving with Graph Visual Question Answering
- LMDrive: Closed-Loop End-to-End Driving with Large Language Models
- Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
- Dolphins: Multimodal Language Model for Driving
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
- A Language Agent for Autonomous Driving
- PaLI-3 Vision Language Models: Smaller, Faster, Stronger
- LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
- Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- GPT-Driver: Learning to Drive with GPT
- MotionLM: Multi-Agent Motion Forecasting as Language Modeling
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes
- SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- VAD: Vectorized Scene Representation for Efficient Autonomous Driving
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Planning-oriented Autonomous Driving
- EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
- Scaling Instruction-Finetuned Language Models
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
- ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning
- Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
- Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
- BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
- UL2: Unifying Language Learning Paradigms
- Symphony: Learning Realistic and Diverse Agents for Autonomous Driving Simulation
- Flamingo: a Visual Language Model for Few-Shot Learning
- BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction
- LoRA: Low-Rank Adaptation of Large Language Models
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset
- Offboard 3D Object Detection from Point Cloud Sequences
- Learning Transferable Visual Models From Natural Language Supervision
- MP3: A Unified Model to Map, Perceive, Predict and Plan
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Curious Case of Neural Text Degeneration
- nuScenes: A multimodal dataset for autonomous driving
- CARLA: An Open Urban Driving Simulator
- End-to-end Driving via Conditional Imitation Learning
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Cited by
Related