Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
2025/07/21 by Jindong Li, Li, Jindong, Yali Fu +13 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2507.22920
openalex publication_date 2025/07/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The rapid advancement of large language models (LLMs) has intensified the need for effective mechanisms to transform continuous multimodal data into discrete representations suitable for language-based processing. Discrete tokenization, with vector quantization (VQ) as a central approach, offers both computational efficiency and compatibility with LLM architectures. Despite its growing importance, there is a lack of a comprehensive survey that systematically examines VQ techniques in the context of LLM-based systems. This work fills this gap by presenting the first structured taxonomy and analysis of discrete tokenization methods designed for LLMs. We categorize 8 representative VQ variants that span classical and modern paradigms and analyze their algorithmic principles, training dynamics, and integration challenges with LLM pipelines. Beyond algorithm-level investigation, we discuss existing research in terms of classical applications without LLMs, LLM-based single-modality systems, and LLM-based multimodal systems, highlighting how quantization strategies influence alignment, reasoning, and generation performance. In addition, we identify key challenges including codebook collapse, unstable gradient estimation, and modality-specific encoding constraints. Finally, we discuss emerging research directions such as dynamic and task-adaptive quantization, unified tokenization frameworks, and biologically inspired codebook learning. This survey bridges the gap between traditional vector quantization and modern LLM applications, serving as a foundational reference for the development of efficient and generalizable multimodal systems. A continuously updated version is available at: https://github.com/jindongli-Ai/LLM-Discrete-Tokenization-Survey.
Citations
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- End-to-End Vision Tokenizer Tuning
- Qwen3 Technical Report
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- Kimi-Audio Technical Report
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
- TVC: Tokenized Video Compression with Ultra-Low Bit Rate
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication
- L3AC: Towards a Lightweight and Lossless Audio Codec
- Universal Item Tokenization for Transferable Generative Recommendation
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
- MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
- Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
- QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model
- HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
- Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
- V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
- SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
- Recent Advances in Discrete Speech Tokens: A Review
- QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
- Multimodal Medical Code Tokenizer
- A Survey of Quantized Graph Representation Learning: Connecting Graph Structures with Large Language Models
- Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
- Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization
- VidTok: A Versatile and Open-Source Video Tokenizer
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
- SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
- Language-Guided Image Tokenization for Generation
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- Scalable Image Tokenization with Index Backpropagation Quantization
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Scaling Transformers for Low-Bitrate High-Quality Speech Coding
- MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
- QARM: Quantitative Alignment Multi-Modal Recommendation at Kuaishou
- GFT: Graph Foundation Model with Transferable Tree Vocabulary
- Image Understanding Makes for A Good Tokenizer for Image Generation
- Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
- LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
- OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
- MotionGlot: A Multi-Embodied Motion Generation Model
- Learning Graph Quantized Tokenizers
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- From Anchors to Answers: A Novel Node Tokenizer for Integrating Graph Structure into Large Language Models
- HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
- IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
- Restructuring Vector Quantization with the Rotation Trick
- LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
- Loong: Generating Minute-level Long Videos with Autoregressive Language Models
- A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
- Emu3: Next-Token Prediction is All You Need
- MIO: A Foundation Model on Multimodal Tokens
- MaskBit: Embedding-free Image Generation via Bit Tokens
- Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
- Moshi: a speech-text foundation model for real-time dialogue
- Learning Multi-Aspect Item Palette: A Semantic Tokenization Framework for Generative Recommendation
- Generative Recommender with End-to-End Learnable Item Tokenization
- Comparing Discrete and Continuous Space LLMs for Speech Recognition
- SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
- UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
- The Llama 3 Herd of Models
- Generative Expressive Conversational Speech Synthesis
- On the Role of Discrete Tokenization in Visual Representation Learning
- MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
- EAGER: Two-Stream Generative Recommender with Behavior-Semantic Collaboration
- Multi-View Empowered Structural Graph Wordification for Language Models
- Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
- ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
- TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
- DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
- Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens
- Image and Video Tokenization with Binary Spherical Quantization
- An Image is Worth 32 Tokens for Reconstruction and Generation
- Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
- Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
- SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
- Node Identifiers: Compact, Discrete Representations for Efficient Graph Learning
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Libra: Building Decoupled Vision System on Large Language Models
- Learnable Item Tokenization for Generative Recommendation
- Vector Quantization for Recommender Systems: A Review and Outlook
- Auto-Encoding Morph-Tokens for Multimodal LLM
- CoST: Contrastive Quantization based Semantic Tokenization for Generative Recommendation
- Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
- SemGrasp: Semantic Grasp Generation via Language Aligned Discretization
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- LLMs are Good Action Recognizers
- Towards Variable and Coordinated Holistic Co-Speech Motion Generation
- GLAD: Improving Latent Graph Generative Modeling with Simple Quantization
- RT-NeRV: Rethinking Hybrid Neural Representations for Video via Residual Tokenization
- HyperVQ: MLR-based Vector Quantization in Hyperbolic Space
- Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
- MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein Embedding
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control
- World Model on Million-Length Video And Language With Blockwise RingAttention
- Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
- StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
- Residual Quantization with Implicit Neural Codebooks
- SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
- WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
- HiHPQ: Hierarchical Hyperbolic Product Quantization for Unsupervised Image Retrieval
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- VideoPoet: A Large Language Model for Zero-Shot Video Generation
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond
- Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation
- TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models
- Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph Representation
- Learning Invariant Molecular Representation in Latent Discrete Space
- Action-Quantized Offline Reinforcement Learning for Robotic Skill Learning
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
- Soft Convex Quantization: Revisiting Vector Quantization with Convex Optimization
- Making LLaMA SEE and Draw with SEED Tokenizer
- Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers
- Finite Scalar Quantization: VQ-VAE Made Simple
- Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
- Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
- Online Clustered Codebook
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Planting a SEED of Vision in Large Language Model
- SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Discrete Graph Auto-Encoder
- High-Fidelity Audio Compression with Improved RVQGAN
- Not All Image Regions Matter: Masked Vector Quantization for Autoregressive Image Generation
- Textually Pretrained Speech Language Models
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector Quantization
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks
- Recommender Systems with Generative Retrieval
- HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
- LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
- Regularized Vector Quantization for Tokenized Image Synthesis
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
- Vector Quantized Wasserstein Auto-Encoder
- Entity-Agnostic Representation Learning for Parameter-Efficient Knowledge Graph Embedding
- Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Muse: Text-To-Image Generation via Masked Generative Transformers
- MAGVIT: Masked Generative Video Transformer
- Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
- Homology-constrained vector quantization entropy regularizer
- MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis
- High Fidelity Neural Audio Compression
- Learning Vector-Quantized Item Representation for Transferable Sequential Recommenders
- Phenaki: Variable Length Video Generation From Open Domain Textual Description
- AudioGen: Textually Guided Audio Generation
- MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation
- BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
- ReFRS: Resource-efficient Federated Recommender System for Dynamic and Diversified User Preferences
- ReFRS: Resource-efficient Federated Recommender System for Dynamic and Diversified User Preferences
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
- OPT: Open Pre-trained Transformer Language Models
- Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
- PaLM: Scaling Language Modeling with Pathways
- Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
- Autoregressive Image Generation using Residual Quantization
- NÜWA-LIP: Language Guided Image Inpainting with Defect-free VQGAN
- MaskGIT: Masked Generative Image Transformer
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- Vector-quantized Image Modeling with Improved VQGAN
- Contrastive Quantization with Code Memory for Unsupervised Image Retrieval
- Self-supervised Product Quantization for Deep Unsupervised Image Retrieval
- SoundStream: An End-to-End Neural Audio Codec
- NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs
- BEiT: BERT Pre-Training of Image Transformers
- CogView: Mastering Text-to-Image Generation via Transformers
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Zero-Shot Text-to-Image Generation
- Taming Transformers for High-Resolution Image Synthesis
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Learning Multi-granular Quantized Embeddings for Large-Vocab Categorical Features in Recommender Systems
- Hierarchical Quantized Autoencoders
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Differentiable Product Quantization for End-to-End Embedding Compression
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- Theory and Experiments on Vector Quantized Autoencoders
- End-to-End Supervised Product Quantization for Image Search and Retrieval
- Neural Discrete Representation Learning
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
- Categorical Reparameterization with Gumbel-Softmax
- Supervised Quantization for Similarity Search
- Transformed Residual Quantization for Approximate Nearest Neighbor Search
- Improved Residual Vector Quantization for High-dimensional Approximate Nearest Neighbor Search
- Stacked Quantizers for Compositional Vector Compression
- Optimized Cartesian K-Means
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Cited by
Related