Visual Representation Alignment for Multimodal Large Language Models
2025/09/09 by Yoon, Heeji, Jung, Jaewoo, Kim, Junwan +10 · 9 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2509.07979
Abstract
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tasks such as object counting or spatial reasoning. We attribute this gap to the prevailing text-only supervision paradigm, which provides only indirect guidance for the visual pathway and often leads MLLMs to discard fine-grained visual details during training. In this paper, we present VIsual Representation ALignment (VIRAL), a simple yet effective regularization strategy that aligns the internal visual representations of MLLMs with those of pre-trained vision foundation models (VFMs). By explicitly enforcing this alignment, VIRAL enables the model not only to retain critical visual details from the input vision encoder but also to complement additional visual knowledge from VFMs, thereby enhancing its ability to reason over complex visual inputs. Our experiments demonstrate consistent improvements across all tasks on widely adopted multimodal benchmarks. Furthermore, we conduct comprehensive ablation studies to validate the key design choices underlying our framework. We believe this simple finding opens up an important direction for the effective integration of visual information in training MLLMs.
Citations
- Lost in Embeddings: Information Loss in Vision-Language Models
- How Visual Representations Map to Language Feature Space in Multimodal LLMs
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
- Perception Encoder: The best visual embeddings are not at the output of the network
- Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
- Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
- Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
- LEO: Boosting Mixture of Vision Encoders for Multimodal Large Language Models
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- FastVLM: Efficient Vision Encoding for Vision Language Models
- Cross-View Completion Models are Zero-shot Correspondence Estimators
- RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- VisionZip: Longer is Better but Not Necessary in Vision Language Models
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
- Cross-modal Information Flow in Multimodal Large Language Models
- What's in the Image? A Deep-Dive into the Vision of Vision Language Models
- Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
- PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
- Reconstructive Visual Instruction Tuning
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- Towards Interpreting Visual Information Processing in Vision-Language Models
- Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
- LLaVA-OneVision: Easy Visual Task Transfer
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Depth Anything V2
- The Platonic Representation Hypothesis
- BRAVE: Broadening the visual encoding of vision-language models
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Unifying Correspondence, Pose and NeRF for Pose-Free Novel View Synthesis from Stereo Pairs
- Honeybee: Locality-enhanced Projector for Multimodal LLM
- AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- What's "up" with vision-language models? Investigating their struggle with spatial reasoning
- Improved Baselines with Visual Instruction Tuning
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence
- What Makes for Good Visual Tokenizers for Large Language Models?
- Evaluating Object Hallucination in Large Vision-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Micrograph segmentations for DDEVD
- Segment Anything
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- CREPE: Can Vision-Language Foundation Models Reason Compositionally?
- Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Cost Aggregation with 4D Convolutional Swin Transformer for Few-Shot Segmentation
- Flamingo: a Visual Language Model for Few-Shot Learning
- CATs++: Boosting Cost Aggregation with Convolutions and Transformers
- Learning Transferable Visual Models From Natural Language Supervision
Cited by
Related