VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
2025/11/28 by Sinan Du, Du, Sinan, Jiahao Guo +18
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis
paper · pdf · doi:10.48550/arxiv.2511.23386
openalex publication_date 2025/11/28 · openalex created_date 2025/12/03 · openalex updated_date 2026/07/28
Abstract
Unifying multimodal understanding, generation and reconstruction representation in a single tokenizer remains a key challenge in building unified models. Previous research predominantly attempts to address this in a dual encoder paradigm, e.g., utilizing the separate encoders for understanding and generation respectively or balancing semantic representations and low-level features with contrastive loss. In this paper, we propose VQRAE, a Vector Quantization version of Representation AutoEncoders, which pioneers the first exploration in unified representation to produce Continuous semantic features for image understanding and Discrete tokens for visual generation within a unified tokenizer. Specifically, we build upon pretrained vision foundation models with a symmetric ViT decoder and adopt a two-stage training strategy: first, it freezes the encoder and learns a high-dimensional semantic VQ codebook with pixel reconstruction objective; then jointly optimizes the encoder with self-distillation constraints. This design enables negligible semantic information for maintaining the ability of multimodal understanding, discrete tokens that are compatible for generation and fine-grained reconstruction. Besides, we identify the intriguing property in quantizing semantic encoders that rely on high-dimensional codebook in contrast to the previous common practice of low-dimensional codebook in image reconstruction. The semantic VQ codebook can achieve a 100% utilization ratio at a dimension of 1536. VQRAE presents competitive performance on several benchmarks of visual understanding, generation and reconstruction with promising scaling property in the autoregressive paradigm for its discrete merits.
Citations
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- Emu3.5: Native Multimodal Models are World Learners
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Diffusion Transformers with Representation Autoencoders
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Reconstruction Alignment Improves Unified Multimodal Models
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Emerging Properties in Unified Multimodal Pretraining
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Qwen3 Technical Report
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Transfer between Modalities with MetaQueries
- ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
- DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
- Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
- SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
- SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- Scalable Image Tokenization with Index Backpropagation Quantization
- MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
- Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- Emu3: Next-Token Prediction is All You Need
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
- ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- Autoregressive Image Generation without Vector Quantization
- An Image is Worth 32 Tokens for Reconstruction and Generation
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
- ChartBench: A Benchmark for Complex Visual Reasoning in Charts
- Generative Multimodal Models are In-Context Learners
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Improved Baselines with Visual Instruction Tuning
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- MMBench: Is Your Multi-modal Model an All-around Player?
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- JourneyDB: A Benchmark for Generative Image Understanding
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Evaluating Object Hallucination in Large Vision-Language Models
- Sigmoid Loss for Language Image Pre-Training
- Scalable Diffusion Models with Transformers
- MAGVIT: Masked Generative Video Transformer
- BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
- MaskGIT: Masked Generative Image Transformer
- High-Resolution Image Synthesis with Latent Diffusion Models
- Emerging Properties in Self-Supervised Vision Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- Zero-Shot Text-to-Image Generation
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
- Taming Transformers for High-Resolution Image Synthesis
- Towards VQA Models That Can Read
- Neural Discrete Representation Learning
- A Diagram Is Worth A Dozen Images
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
Cited by
Related