Let ViT Speak: Generative Language-Image Pre-training
2026/05/01 by Yan Fang, Mengcheng Lan, Zilong Huang +7 · 1 voice
Computer Science · #Domain Adaptation and Few-Shot Learning #Encoder #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Key (lock) #Language acquisition #Language model #Multimodal Machine Learning Applications #Natural language #Transformer #cs.CV
paper · pdf · open access · doi:10.48550/arxiv.2605.00809
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2026/05/01 · arxiv published 2026/05/01 · openalex created_date 2026/05/05 · arxiv updated 2026/06/09 · openalex updated_date 2026/07/28
Abstract
In this paper, we present Generative Language-Image Pre-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers (ViTs) designed for multimodal large language models (MLLMs). To better align vision encoders with the autoregressive nature of LLMs, GenLIP trains a ViT to predict language tokens directly from visual tokens using a standard language modeling objective, without contrastive batch construction or an additional text decoder. This design offers three key advantages: (1) Simplicity: a single transformer jointly models visual and textual tokens; (2) Scalability: it scales effectively with both data and model size; and (3) Performance: it achieves competitive or superior results across diverse multimodal benchmarks. Trained on 8B samples from Recap-DataComp-1B, GenLIP matches or surpasses strong baselines despite using substantially less pretraining data. After continued pretraining on multi-resolution images at native aspect ratios, GenLIP further improves on detail-sensitive tasks such as OCR and chart understanding, making it a strong foundation for vision encoders in MLLMs.
Citations
- Qwen3-VL Technical Report
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Meta CLIP 2: A Worldwide Scaling Recipe
- Region-based Cluster Discrimination for Visual Representation Learning
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- Qwen3 Technical Report
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- Kimi-VL Technical Report
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
- Systematic Outliers in Large Language Models
- CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
- Multimodal Autoregressive Pre-training of Large Vision Encoders
- Classification Done Right for Vision-Language Pre-Training
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LLaVA-OneVision: Easy Visual Task Transfer
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
- PaliGemma: A versatile 3B VLM for transfer
- SOLO: A Single Transformer for Scalable Vision-Language Modeling
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Unveiling Encoder-Free Vision-Language Models
- What If We Recaption Billions of Web Images with LLaMA-3?
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- DreamLIP: Language-Image Pre-training with Long Captions
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Generative Multimodal Models are In-Context Learners
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- VeCLIP: Improving CLIP Training via Visual-enriched Captions
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Efficient Streaming Language Models with Attention Sinks
- Vision Transformers Need Registers
- Demystifying CLIP Data
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- ALIP: Adaptive Language-Image Pre-training with Synthetic Caption
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Image Captioners Are Scalable Vision Learners Too
- Improving CLIP Training with Language Rewrites
- DataComp: In search of the next generation of multimodal datasets
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Unifying Vision-Language Representation Space with Single-tower Transformer
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- VL-BEiT: Generative Vision-Language Pretraining
- GIT: A Generative Image-to-text Transformer for Vision and Language
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- InfographicVQA
- Learning Transferable Visual Models From Natural Language Supervision
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- DocVQA: A Dataset for VQA on Document Images
- TextCaps: a Dataset for Image Captioning with Reading Comprehension
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Towards VQA Models That Can Read
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- A Diagram Is Worth A Dozen Images
- Generation and Comprehension of Unambiguous Object Descriptions
Discussions
Related