Transfer between Modalities with MetaQueries
2025/04/08 by Xichen Pan, Satya Narayan Shukla, Pan, Xichen +24 · 7 voices · 84 citations
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
paper · pdf · doi:10.48550/arxiv.2504.06256
Abstract
Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often demands complex training recipes and careful data balancing. We introduce MetaQueries, a set of learnable queries that act as an efficient interface between autoregressive multimodal LLMs (MLLMs) and diffusion models. MetaQueries connects the MLLM's latents to the diffusion decoder, enabling knowledge-augmented image generation by leveraging the MLLM's deep understanding and reasoning capabilities. Our method simplifies training, requiring only paired image-caption data and standard diffusion objectives. Notably, this transfer is effective even when the MLLM backbone remains frozen, thereby preserving its state-of-the-art multimodal understanding capabilities while achieving strong generative performance. Additionally, our method is flexible and can be easily instruction-tuned for advanced applications such as image editing and subject-driven generation.
Citations
Cited by
- SuperFlow: Training Flow Matching Models with RL on the Fly
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- Conditioning Residuals for Diffusion Models via Representation Feedback
- HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
- ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- LongCat-Image Technical Report
- ThinkGen: Generalized Thinking for Visual Generation
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- What Happens Next? Next Scene Prediction with a Unified Video Model
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- MolSculpt: Sculpting 3D Molecular Geometries from Chemical Syntax
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- UniLight: A Unified Representation for Lighting
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Plan-X: Instruct Video Generation via Semantic Planning
- UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- Optimizing Input of Denoising Score Matching is Biased Towards Higher Score Norm
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- Diffusion Transformers with Representation Autoencoders
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Unified Multimodal Model as Auto-Encoder
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Reconstruction Alignment Improves Unified Multimodal Models
- Interleaving Reasoning for Better Text-to-Image Generation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- SpotEdit: Evaluating Visually-Guided Image Editing Methods
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Discussions
- Transfer between Modalities with MetaQueries [hn, 25 points, 12 comments]
- Мое предсказание после генерации изображений GPT-4 #ai #gpt #news [bsky, 0 points, 0 comments]
- Transfer between Modalities with MetaQueries https://arxiv.org/abs/2504.06256 (https://news.ycombinator.com/item?id=43667963) [bsky, 0 points, 0 comments]
- My prediction after GPT-4o image generation #ai #gpt #news [bsky, 0 points, 0 comments]
- My prediction after GPT-4o image generation #HackerNews https://arxiv.org/abs/2504.06256 [bsky, 0 points, 0 comments]
- My prediction after GPT-4o image generation https://arxiv.org/abs/2504.06256 [bsky, 0 points, 0 comments]
- My prediction after GPT-4o image generation [bsky, 0 points, 0 comments]
Related