2024/09/06 by Jianbiao Mei, Mei, Jianbiao, Hu, Tao +15 · 14 citations
Computer Science · #Advanced Vision and Imaging #Artificial intelligence #Autoregressive model #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #Computer graphics (images) #Computer science #Computer vision #Econometrics #Economics #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Motion (physics)
paper · pdf · doi:10.48550/arxiv.2409.04003
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/09/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Recent advances in diffusion models have improved controllable streetscape generation and supported downstream perception and planning tasks. However, challenges remain in accurately modeling driving scenes and generating long videos. To alleviate these issues, we propose DreamForge, an advanced diffusion-based autoregressive video generation model tailored for 3D-controllable long-term generation. To enhance the lane and foreground generation, we introduce perspective guidance and integrate object-wise position encoding to incorporate local 3D correlation and improve foreground object modeling. We also propose motion-aware temporal attention to capture motion cues and appearance changes in videos. By leveraging motion frames and an autoregressive generation paradigm,we can autoregressively generate long videos (over 200 frames) using a model trained in short sequences, achieving superior quality compared to the baseline in 16-frame video evaluations. Finally, we integrate our method with the realistic simulator DriveArena to provide more reliable open-loop and closed-loop evaluations for vision-based driving agents. Project Page: https://pjlab-adg.github.io/DriveArena/dreamforge.