UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
2025/06/25 by Yanzhe Chen, Huasong Zhong, Chen, Yanzhe +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimedia (cs.MM) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.20214
openalex publication_date 2025/06/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based methods either rely on small vocabularies (~16K entries) that lack fine-grained semantics or naively scale up, resulting in low token utilization and unstable training. We propose UniCode2, a cascaded codebook framework enabling large-scale, semantically aligned, and stable visual tokenization. By clustering millions of SigLIP sequence embeddings, we build a 500K-entry codebook that preserves vision-language alignment while expanding capacity. Stability is ensured via a cascaded design: a frozen codebook anchors the embedding space, and a trainable codebook refines task-specific semantics. This decoupling promotes high utilization and robust learning. Moreover, the alignment of our visual tokens with textual semantics enables seamless integration with pretrained diffusion decoders, supporting high-quality visual synthesis with minimal adaptation. UniCode2 delivers strong performance across diverse benchmarks, demonstrating the viability of scaling visual token spaces without sacrificing stability, semantics, or modularity.
Citations
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
- DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
- OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
- SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
- Qwen2.5-VL Technical Report
- QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
- Qwen2.5 Technical Report
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
- Factorized Visual Tokenization and Generation
- Image Understanding Makes for A Good Tokenizer for Image Generation
- Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Emu3: Next-Token Prediction is All You Need
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- OmniGen: Unified Image Generation
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- LLaVA-OneVision: Easy Visual Task Transfer
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- World Model on Million-Length Video And Language With Blockwise RingAttention
- Generative Multimodal Models are In-Context Learners
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- Improved Baselines with Visual Instruction Tuning
- Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers
- Qwen Technical Report
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- NExT-GPT: Any-to-Any Multimodal LLM
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Online Clustered Codebook
- MMBench: Is Your Multi-modal Model an All-around Player?
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- JourneyDB: A Benchmark for Generative Image Understanding
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- EVA-02: A visual representation for neon genesis
- GPT-4 Technical Report
- A Survey on Self-supervised Learning: Algorithms, Applications, and Future Trends
- EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
- BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
- Autoregressive Image Generation using Residual Quantization
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Vector-quantized Image Modeling with Improved VQGAN
- BEiT: BERT Pre-Training of Image Transformers
- Taming Transformers for High-Resolution Image Synthesis
- Towards VQA Models That Can Read
- Neural Discrete Representation Learning
- A Diagram Is Worth A Dozen Images
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
- Scalable Image Tokenization with Index Backpropagation Quantization
Related