vix.ing · top · new · best · stats

Doe-1: Closed-Loop Autonomous Driving with Large World Model

2024/12/12 by Wenzhao Zheng, Zheng, Wenzhao, Zetian Xia +9 · 1 voice · 19 citations
Computer Science · Decision Sciences · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Artificial intelligence #Autonomous Vehicle Technology and Safety #Closed loop #Combinatorics #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Control (management) #Control engineering #Control theory (sociology) #Engineering #FOS: Computer and information sciences #Loop (graph theory) #Machine Learning (cs.LG) #Mathematics #Simulation Techniques and Applications #Traffic Prediction and Management Techniques #cs.AI #cs.CV #cs.LG

paper · pdf · doi:10.48550/arxiv.2412.09627

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/12/12 · arxiv published 2024/12/12 · arxiv updated 2024/12/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

End-to-end autonomous driving has received increasing attention due to its potential to learn from large amounts of data. However, most existing methods are still open-loop and suffer from weak scalability, lack of high-order interactions, and inefficient decision-making. In this paper, we explore a closed-loop framework for autonomous driving and propose a large Driving wOrld modEl (Doe-1) for unified perception, prediction, and planning. We formulate autonomous driving as a next-token generation problem and use multi-modal tokens to accomplish different tasks. Specifically, we use free-form texts (i.e., scene descriptions) for perception and generate future predictions directly in the RGB space with image tokens. For planning, we employ a position-aware tokenizer to effectively encode action into discrete tokens. We train a multi-modal transformer to autoregressively generate perception, prediction, and planning tokens in an end-to-end and unified manner. Experiments on the widely used nuScenes dataset demonstrate the effectiveness of Doe-1 in various tasks including visual question-answering, action-conditioned video generation, and motion planning. Code: https://github.com/wzzheng/Doe.

Cited by

Discussions

Related