Chameleon: Mixed-Modal Early-Fusion Foundation Models
2024/05/16 by Chameleon Team · 9 voices · 240 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #cs.CL
paper · pdf · doi:10.48550/arxiv.2405.09818
openalex publication_date 2024/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
We present Chameleon, a family of early-fusion token-based mixed-modal models capable of understanding and generating images and text in any arbitrary sequence. We outline a stable training approach from inception, an alignment recipe, and an architectural parameterization tailored for the early-fusion, token-based, mixed-modal setting. The models are evaluated on a comprehensive range of tasks, including visual question answering, image captioning, text generation, image generation, and long-form mixed modal generation. Chameleon demonstrates broad and general capabilities, including state-of-the-art performance in image captioning tasks, outperforms Llama-2 in text-only tasks while being competitive with models such as Mixtral 8x7B and Gemini-Pro, and performs non-trivial image generation, all in a single model. It also matches or exceeds the performance of much larger models, including Gemini Pro and GPT-4V, according to human judgments on a new long-form mixed-modal generation evaluation, where either the prompt or outputs contain mixed sequences of both images and text. Chameleon marks a significant step forward in a unified modeling of full multimodal documents.
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
- T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
- STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- On the Theoretical Limitations of Embedding-Based Retrieval
- WorldVLA: Towards Autoregressive Action World Model
- Transfer between Modalities with MetaQueries
- Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- VGGT-Ω
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Large Video Planner Enables Generalizable Robot Control
- FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- Motus: A Unified Latent Action World Model
- What Happens Next? Next Scene Prediction with a Unified Video Model
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Grounding Everything in Tokens for Multimodal Large Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- PhotoFramer: Multi-modal Image Composition Instruction
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- Visual Generation Tuning
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Spanning Tree Autoregressive Visual Generation
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
- Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Emu3.5: Native Multimodal Models are World Learners
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- PairUni: Pairwise Training for Unified Multimodal Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- End-to-End Multi-Modal Diffusion Mamba
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Task-Aware Resolution Optimization for Visual Large Language Models
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- Latent Visual Reasoning
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- Planning with Unified Multimodal Models
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
- AToken: A Unified Tokenizer for Vision
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Unified Multimodal Model as Auto-Encoder
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- Reconstruction Alignment Improves Unified Multimodal Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Interleaving Reasoning for Better Text-to-Image Generation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- A Survey on Training-free Alignment of Large Language Models
- Grouped Speculative Decoding for Autoregressive Image Generation
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Multimodal learning with next-token prediction for large multimodal models
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- The Promise of RL for Autoregressive Image Editing
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Meta CLIP 2: A Worldwide Scaling Recipe
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- FedVLM: Scalable Personalized Vision-Language Models through Federated Learning
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- Automating Steering for Safe Multimodal Large Language Models
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- Emergence of Hierarchical Emotion Organization in Large Language Models
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
- FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- MotionGPT3: Human Motion as a Second Modality
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- Diffusion model [wikipedia]
Discussions
- Chameleon: Meta’s New Multi-Modal LLM [hn, 304 points, 40 comments]
- What are some must-read multimodal generation papers? I am looking for vision-language models (preferably jointly trained from scratch). Some examples - Chameleon - arxiv.org/abs/2405.09818 Transfus [bsky, 27 points, 1 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 6 points, 0 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 4 points, 0 comments]
- The hope for a lot of these models were modalities are tightly integrated (often called "natively multi-modal" or "early fusion") is that there is synergy between modalities but one might also find th [bsky, 3 points, 2 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [lobsters, 1 points, 0 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 1 points, 0 comments]
- Chameleon: Meta's New Multi-Modal LLM (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- New #multimodal #LLM from Meta's #chameleon team! arxiv.org/abs/2405.09818 #academics #AI #ML [bsky, 0 points, 0 comments]
Related