AToken: A Unified Tokenizer for Vision
2025/09/17 by Lu, Jiasen, Song, Liangchen, Xu, Mingze +5 · 8 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2509.14476
Abstract
We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existing tokenizers that specialize in either reconstruction or understanding for single modalities, AToken encodes these diverse visual inputs into a shared 4D latent space, unifying both tasks and modalities in a single framework. Specifically, we introduce a pure transformer architecture with 4D rotary position embeddings to process visual inputs of arbitrary resolutions and temporal durations. To ensure stable training, we introduce an adversarial-free training objective that combines perceptual and Gram matrix losses, achieving state-of-the-art reconstruction quality. By employing a progressive training curriculum, AToken gradually expands from single images, videos, and 3D, and supports both continuous and discrete latent tokens. AToken achieves 0.21 rFID with 82.2% ImageNet accuracy for images, 3.01 rFVD with 40.2% MSRVTT retrieval for videos, and 28.28 PSNR with 90.9% classification accuracy for 3D.. In downstream applications, AToken enables both visual generation tasks (e.g., image generation with continuous and discrete tokens, text-to-video generation, image-to-3D synthesis) and understanding tasks (e.g., multimodal LLMs), achieving competitive performance across all benchmarks. These results shed light on the next-generation multimodal AI systems built upon unified visual tokenization.
Citations
- Show-o2: Improved Native Unified Multimodal Models
- Emerging Properties in Unified Multimodal Pretraining
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Perception Encoder: The best visual embeddings are not at the output of the network
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Wan: Open and Advanced Large-Scale Video Generative Models
- SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
- Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
- Open-Sora 2.0: Training a Commercial-Level Video Generation Model in 200k
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Masked Autoencoders Are Effective Tokenizers for Diffusion Models
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
- Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
- Scaling 4D Representations
- SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
- Apollo: An Exploration of Video Understanding in Large Multimodal Models
- Language-Guided Image Tokenization for Generation
- LinVT: Empower Your Image-level Large Language Model to Understand Videos
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Structured 3D Latents for Scalable and Versatile 3D Generation
- XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
- LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Movie Gen: A Cast of Media Foundation Models
- ElasticTok: Adaptive Tokenization for Image and Video
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- ImageFolder: Autoregressive Image Generation with Folded Tokens
- MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
- OctFusion: Octree-based Diffusion Models for 3D Shape Generation
- Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- LLaVA-OneVision: Easy Visual Task Transfer
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
- LVBench: An Extreme Long Video Understanding Benchmark
- An Image is Worth 32 Tokens for Reconstruction and Generation
- Towards Semantic Equivalence of Tokenization in Multimodal LLM
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- View Selection for 3D Captioning via Diffusion Ranking
- LongVLM: Efficient Long Video Understanding via Large Language Models
- GVGEN: Text-to-3D Generation with Volumetric Representation
- LN3DIFF++: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- VideoPrism: A Foundational Visual Encoder for Video Understanding
- SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- 4M: Massively Multimodal Masked Modeling
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Demystifying CLIP Data
- Finite Scalar Quantization: VQ-VAE Made Simple
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
- Objaverse-XL: A Universe of 10M+ 3D Objects
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- VideoLLM: Modeling Video Sequence with Large Language Models
- A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
- Shap-E: Generating Conditional 3D Implicit Functions
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- GPT-4 Technical Report
- 3DGen: Triplane Latent Diffusion for Textured Mesh Generation
- LLaMA: Open and Efficient Foundation Language Models
- Scalable Diffusion Models with Transformers
- Point-E: A System for Generating 3D Point Clouds from Complex Prompts
- Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion
- MAGVIT: Masked Generative Video Transformer
- InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- 3D Neural Field Generation using Triplane Diffusion
- Phenaki: Variable Length Video Generation From Open Domain Textual Description
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Neural Wavelet-domain Diffusion for 3D Shape Generation
- MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation
- Clover: Towards A Unified Video-Language Alignment and Fusion Model
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
- PaLM: Scaling Language Modeling with Pathways
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- Autoregressive Image Generation using Residual Quantization
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
- MERLOT Reserve: Neural Script Knowledge through Vision and Language and\n Sound
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Masked Autoencoders Are Scalable Vision Learners
- Vector-quantized Image Modeling with Improved VQGAN
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- BEiT: BERT Pre-Training of Image Transformers
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- VideoGPT: Video Generation using VQ-VAE and Transformers
- CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- Broaden Your Views for Self-Supervised Video Learning
- Diffusion Probabilistic Models for 3D Point Cloud Generation
- Learning Transferable Visual Models From Natural Language Supervision
- Using Shape to Categorize: Low-Shot Learning with an Explicit Shape Bias
- Taming Transformers for High-Resolution Image Synthesis
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Spatiotemporal Contrastive Video Representation Learning
- A Simple Framework for Contrastive Learning of Visual Representations
- Towards VQA Models That Can Read
- Occupancy Networks: Learning 3D Reconstruction in Function Space
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Neural Discrete Representation Learning
- The 2017 DAVIS Challenge on Video Object Segmentation
- A Diagram Is Worth A Dozen Images
- Neural Machine Translation of Rare Words with Subword Units
- Improving neural networks by preventing co-adaptation of feature detectors
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- MLVU: Benchmarking Multi-task Long Video Understanding
Cited by
Related