Video models are zero-shot learners and reasoners
2025/09/24 by Thaddäus Wiedemer, Yuxuan Li, Wiedemer, Thaddäus +15 · 7 voices · 55 citations
#cs.LG #cs.AI #cs.CV #cs.RO
paper · pdf · doi:10.48550/arxiv.2509.20328
Abstract
The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. This transformation emerged from simple primitives: large, generative models trained on web-scale data. Curiously, the same primitives apply to today's generative video models. Could video models be on a trajectory towards general-purpose vision understanding, much like LLMs developed general-purpose language understanding? We demonstrate that Veo 3 can solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and more. These abilities to perceive, model, and manipulate the visual world enable early forms of visual reasoning like maze and symmetry solving. Veo's emergent zero-shot capabilities indicate that video models are on a path to becoming unified, generalist vision foundation models.
Citations
Cited by
- Video Understanding: From Geometry and Semantics to Unified Models
- Visual prompt engineering for video models
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
- Kling-Omni Technical Report
- In Pursuit of Pixel Supervision for Visual Pre-training
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- Unified Video Editing with Temporal Reasoner
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
- Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
- Are Image-to-Video Models Good Zero-Shot Image Editors?
- In-Video Instructions: Visual Signals as Generative Control
- Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
- Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
- Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
- First Frame Is the Place to Go for Video Content Customization
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- Walk through Paintings: Egocentric World Models from Internet Priors
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- Learning to Generate Rigid Body Interactions with Video Diffusion Models
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
Discussions
- Paper: arxiv.org/pdf/2509.20328 [bsky, 15 points, 1 comments]
- 🔥Veo 3 has emergent zero-shot learning and reasoning capabilities! This multitalented model can do a huge range of interesting tasks. It understands physical properties, can manipulate objects, and c [bsky, 9 points, 1 comments]
- Video models are zero-shot learners and reasoners [hn, 2 points, 0 comments]
- arxiv.org/pdf/2509.20328 [bsky, 2 points, 0 comments]
- Video models are zero-shot learners and reasoners www.arxiv.org/abs/2509.20328 👀 [bsky, 1 points, 0 comments]
- DeepMind再掀视频AI革命!最新论文提出“帧链(CoF)”概念,给视频模型装上“视觉大脑”,复刻语言模型的链式思维能力 其核心是逐帧推理:每一帧生成都基于前序内容,像导演拍电影般规划,让视频“动得有理、看得真实”,60秒生成任务逻辑错误率直降47%。以Veo 3为代表的模型已展现通用视觉能力,迷宫推理、空间建模样样精通,零样本搞定“看-想”全流程 DeepMind预言,视频“通才”将取代“专 [bsky, 1 points, 0 comments]
- arxiv.org/abs/2509.20328 [bsky, 1 points, 0 comments]
Related