Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
2026/07/27 by Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20
#cs.CV #cs.CL
paper · pdf
Abstract
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.
Citations
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
- Streaming Video Instruction Tuning
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Towards Cross-View Point Correspondence in Vision-Language Models
- Qwen3-VL Technical Report
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
- FineVision: Open Data Is All You Need
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- DINOv3
- Meta CLIP 2: A Worldwide Scaling Recipe
- Ovis-U1 Technical Report
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- Qwen3 Technical Report
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Kimi-VL Technical Report
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Towards Practical Real-Time Neural Video Compression
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
- SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
- Multimodal Autoregressive Pre-training of Large Vision Encoders
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- LLaVA-OneVision: Easy Visual Task Transfer
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- VISA: Reasoning Video Object Segmentation via Large Language Models
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
- Long Context Transfer from Language to Vision
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- VideoLLM-online: Online Video Large Language Model for Streaming Video
- LVBench: An Extreme Long Video Understanding Benchmark
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
- MLVU: Benchmarking Multi-task Long Video Understanding
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- TempCompass: Do Video LLMs Really Understand Videos?
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- Document Understanding Dataset and Evaluation (DUDE)
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- VideoChat: Chat-Centric Video Understanding
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Hierarchical multimodal transformers for Multi-Page DocVQA
- Token Merging: Your ViT But Faster
- Flamingo: a Visual Language Model for Few-Shot Learning
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Video Swin Transformer
- DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- ViViT: A Video Vision Transformer
- Learning Transferable Visual Models From Natural Language Supervision
- Is Space-Time Attention All You Need for Video Understanding?
- WebSRC: A Dataset for Web-Based Structural Reading Comprehension
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- DocVQA: A Dataset for VQA on Document Images
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
- Towards VQA Models That Can Read
- Video Object Segmentation with Language Referring Expressions
- Compressed Video Action Recognition
- The Kinetics Human Action Video Dataset
- TALL: Temporal Activity Localization via Language Query
- A Diagram Is Worth A Dozen Images
- ImageNet Large Scale Visual Recognition Challenge
- Describing Textures in the Wild
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- Separate visual pathways for perception and action
Related