ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
2024/03/08 by Xiwei Hu, Rui Wang, Hu, Xiwei +9 · 200 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2403.05135
Abstract
Diffusion models have demonstrated remarkable performance in the domain of text-to-image generation. However, most widely used models still employ CLIP as their text encoder, which constrains their ability to comprehend dense prompts, encompassing multiple objects, detailed attributes, complex relationships, long-text alignment, etc. In this paper, we introduce an Efficient Large Language Model Adapter, termed ELLA, which equips text-to-image diffusion models with powerful Large Language Models (LLM) to enhance text alignment without training of either U-Net or LLM. To seamlessly bridge two pre-trained models, we investigate a range of semantic alignment connector designs and propose a novel module, the Timestep-Aware Semantic Connector (TSC), which dynamically extracts timestep-dependent conditions from LLM. Our approach adapts semantic features at different stages of the denoising process, assisting diffusion models in interpreting lengthy and intricate prompts over sampling timesteps. Additionally, ELLA can be readily incorporated with community models and tools to improve their prompt-following capabilities. To assess text-to-image models in dense prompt following, we introduce Dense Prompt Graph Benchmark (DPG-Bench), a challenging benchmark consisting of 1K dense prompts. Extensive experiments demonstrate the superiority of ELLA in dense prompt following compared to state-of-the-art methods, particularly in multiple object compositions involving diverse attributes and relationships.
Cited by
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- LongCat-Image Technical Report
- ThinkGen: Generalized Thinking for Visual Generation
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling
- Normalizing Trajectory Models
- InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
- Text-Conditioned Background Generation for Editable Multi-Layer Documents
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- PixelArena: A benchmark for Pixel-Precision Visual Intelligence
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Learning by Analogy: A Causal Framework for Composition Generalization
- LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Glance: Accelerating Diffusion Models with 1 Sample
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Ovis-Image Technical Report
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Distribution Matching Distillation Meets Reinforcement Learning
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
- Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Emu3.5: Native Multimodal Models are World Learners
- LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- L2P: Unlocking Latent Potential for Pixel Generation
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
- Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation
- Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Asymmetric Flow Models
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
- GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- CO3: Contrasting Concepts Compose Better
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Structured Information for Improving Spatial Relationships in Text-to-Image Generation
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Unified Multimodal Model as Auto-Encoder
- MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- Reinforcement Learning for Large Model: A Survey
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Multimodal learning with next-token prediction for large multimodal models
- Elastic Diffusion Transformer
- VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- Qwen-Image Technical Report
- LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
- PixNerd: Pixel Neural Field Diffusion
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
- LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
- TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
- Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
- Pyramidal Patchification Flow for Visual Generation
- Ovis-U1 Technical Report
- XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Show-o2: Improved Native Unified Multimodal Models
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
- Prompt-Guided Latent Diffusion with Predictive Class Conditioning for 3D Prostate MRI Generation
- FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
- Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
- ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
- Image Generation from Contextually-Contradictory Prompts
- Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
- ComposeAnything: Composite Object Priors for Text-to-Image Generation
- ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filtering
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Thinking with Generated Images
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
- HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
- SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
- MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Towards Self-Improvement of Diffusion Models via Group Preference Optimization
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Taming Outlier Tokens in Diffusion Transformers
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Large Language Models are Universal Reasoners for Visual Generation
- Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
- TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
- Continuous Adversarial Flow Models
- Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
- Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
- Energy-Guided Flow Matching
- MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- Taming Consistency Distillation for Accelerated Human Image Animation
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- PixelFlow: Pixel-Space Generative Models with Flow
- Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
Related