Chameleon: Mixed-Modal Early-Fusion Foundation Models
2024/05/16 by Chameleon Team · 9 voices · 368 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Archaeology #Computer science #Foundation (evidence) #Fusion #Generative Adversarial Networks and Image Synthesis #Geography #Linguistics #Materials science #Modal #Multimodal Machine Learning Applications #Philosophy #cs.CL
paper · pdf · doi:10.48550/arxiv.2405.09818
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
We present Chameleon, a family of early-fusion token-based mixed-modal models capable of understanding and generating images and text in any arbitrary sequence. We outline a stable training approach from inception, an alignment recipe, and an architectural parameterization tailored for the early-fusion, token-based, mixed-modal setting. The models are evaluated on a comprehensive range of tasks, including visual question answering, image captioning, text generation, image generation, and long-form mixed modal generation. Chameleon demonstrates broad and general capabilities, including state-of-the-art performance in image captioning tasks, outperforms Llama-2 in text-only tasks while being competitive with models such as Mixtral 8x7B and Gemini-Pro, and performs non-trivial image generation, all in a single model. It also matches or exceeds the performance of much larger models, including Gemini Pro and GPT-4V, according to human judgments on a new long-form mixed-modal generation evaluation, where either the prompt or outputs contain mixed sequences of both images and text. Chameleon marks a significant step forward in a unified modeling of full multimodal documents.
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
- T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
- STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- On the Theoretical Limitations of Embedding-Based Retrieval
- WorldVLA: Towards Autoregressive Action World Model
- Transfer between Modalities with MetaQueries
- Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- VGGT-Ω
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Large Video Planner Enables Generalizable Robot Control
- FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- Motus: A Unified Latent Action World Model
- What Happens Next? Next Scene Prediction with a Unified Video Model
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Grounding Everything in Tokens for Multimodal Large Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- PhotoFramer: Multi-modal Image Composition Instruction
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- Visual Generation Tuning
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Spanning Tree Autoregressive Visual Generation
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
- Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Emu3.5: Native Multimodal Models are World Learners
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- PairUni: Pairwise Training for Unified Multimodal Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- End-to-End Multi-Modal Diffusion Mamba
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Task-Aware Resolution Optimization for Visual Large Language Models
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- Latent Visual Reasoning
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- Planning with Unified Multimodal Models
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
- AToken: A Unified Tokenizer for Vision
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Unified Multimodal Model as Auto-Encoder
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- Reconstruction Alignment Improves Unified Multimodal Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Interleaving Reasoning for Better Text-to-Image Generation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- Vision Generalist Model: A Survey
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- A Survey on Training-free Alignment of Large Language Models
- Grouped Speculative Decoding for Autoregressive Image Generation
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Multimodal learning with next-token prediction for large multimodal models
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- The Promise of RL for Autoregressive Image Editing
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Meta CLIP 2: A Worldwide Scaling Recipe
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- FedVLM: Scalable Personalized Vision-Language Models through Federated Learning
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- Automating Steering for Safe Multimodal Large Language Models
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- Emergence of Hierarchical Emotion Organization in Large Language Models
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
- FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- Demystifying Video Reasoning
- MotionGPT3: Human Motion as a Second Modality
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- TAMMs: Temporal-Aware Multimodal Model for Satellite Image Change Understanding and Forecasting
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- Watermarking Autoregressive Image Generation
- Aligning Text, Images, and 3D Structure Token-by-Token
- Show-o2: Improved Native Unified Multimodal Models
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
- Fake it till You Make it: Reward Modeling as Discriminative Prediction
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- A Watermark for Auto-Regressive Image Generation Models
- SpectralAR: Spectral Autoregressive Visual Generation
- TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
- PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
- Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
- MiMo-VL Technical Report
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- MLaGA: Multimodal Large Language and Graph Assistant
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
- Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Benchmarking Foundation Models for Zero-Shot Biometric Tasks
- Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- REOrdering Patches Improves Vision Models
- D-AR: Diffusion via Autoregressive Models
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Thinking with Generated Images
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- STRICT: Stress Test of Rendering Images Containing Text
- MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- R-Genie: Reasoning-Guided Generative Image Editing
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- ChemMLLM: Chemical Multimodal Large Language Model
- MMaDA: Multimodal Large Diffusion Language Models
- IA-T2I: Internet-Augmented Text-to-Image Generation
- One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
- Emerging Properties in Unified Multimodal Pretraining
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- Video-GPT via Next Clip Diffusion
- Where did the ambiguity go? Examining how multimodal models interpret polysemous words
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Visual Instruction Tuning with Chain of Region-of-Interest
- The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- Vision as Unified Multimodal Generation
- A Survey of Interactive Generative Video
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- The Design Space of Tri-Modal Masked Diffusion Models
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- ETCHR: Editing To Clarify and Harness Reasoning
- Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation
- X-Fusion: Introducing New Modality to Frozen Large Language Models
- YoChameleon: Personalized Vision and Language Generation
- Learning Streaming Video Representation via Multitask Training
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Let ViT Speak: Generative Language-Image Pre-training
- Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Personalized Text-to-Image Generation with Auto-Regressive Models
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions
- A Survey on Efficient Vision-Language Models
- FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
- Scaling Laws for Native Multimodal Models
- Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
- OmniCaptioner: One Captioner to Rule Them All
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
- Diffusion model [wikipedia]
Discussions
- Chameleon: Meta’s New Multi-Modal LLM [hn, 304 points, 40 comments]
- What are some must-read multimodal generation papers? I am looking for vision-language models (preferably jointly trained from scratch). Some examples - Chameleon - arxiv.org/abs/2405.09818 Transfus [bsky, 27 points, 1 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 6 points, 0 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 4 points, 0 comments]
- The hope for a lot of these models were modalities are tightly integrated (often called "natively multi-modal" or "early fusion") is that there is synergy between modalities but one might also find th [bsky, 3 points, 2 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [lobsters, 1 points, 0 comments]
- Chameleon: Mixed-Modal Early-Fusion Foundation Models [hn, 1 points, 0 comments]
- Chameleon: Meta's New Multi-Modal LLM (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- New #multimodal #LLM from Meta's #chameleon team! arxiv.org/abs/2405.09818 #academics #AI #ML [bsky, 0 points, 0 comments]
Related