Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
2026/08/05 by Junlin Han, Shengbang Tong, David Fan +4
Computer Science · #cs.CV #cs.LG #cs.MM
paper · pdf
Project page: https://junlinhan.github.io/projects/physics_of_mm_pretrain/
arxiv created 2026/08/06 · arxiv updated 2026/08/07
Abstract
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data "complexity" largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.
Citations
- Scaling Native Multimodal Pre-Training From Scratch
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- ERNIE 5.0 Technical Report
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- Qwen3-VL Technical Report
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Diffusion Transformers with Representation Autoencoders
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Words That Make Language Models Perceive
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Qwen-Image Technical Report
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- Proof of a perfect platonic representation hypothesis
- Ovis-U1 Technical Report
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- Emerging Properties in Unified Multimodal Pretraining
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Scaling Laws for Native Multimodal Models
- Transfer between Modalities with MetaQueries
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Scaling Language-Free Visual Representation Learning
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
- QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
- Probing Visual Language Priors in VLMs
- LMFusion: Adapting Pretrained Language Models for Multimodal Generation
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
- JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
- Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
- Image Understanding Makes for A Good Tokenizer for Image Generation
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Aria: An Open Multimodal Native Mixture-of-Experts Model
- Emu3: Next-Token Prediction is All You Need
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- The Llama 3 Herd of Models
- MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- The Platonic Representation Hypothesis
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- Grounded language acquisition through the eyes and ears of a single child
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- Generative Multimodal Models are In-Context Learners
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- Investigating the Catastrophic Forgetting in Multimodal Large Language Models
- Planting a SEED of Vision in Large Language Model
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Learning high-level visual representations from a child's perspective without strong inductive biases
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- LLaMA: Open and Efficient Foundation Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- Flamingo: a Visual Language Model for Few-Shot Learning
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Autoregressive Image Generation using Residual Quantization
- CM3: A Causal Masked Multimodal Model of the Internet
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- DocVQA: A Dataset for VQA on Document Images
- AI2D-RST: A multimodal corpus of 1000 primary school science diagrams
- PIQA: Reasoning about Physical Commonsense in Natural Language
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Towards VQA Models That Can Read
- CoQA: A Conversational Question Answering Challenge
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Neural Discrete Representation Learning
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for\n Reading Comprehension
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Semantic Understanding of Scenes through the ADE20K Dataset
- Semantic Understanding of Scenes Through the ADE20K Dataset
- Microsoft COCO: Common Objects in Context
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- The Development of Embodied Cognition: Six Lessons from Babies