Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
2025/12/23 by Byung-Kwan Lee, Lee, Byung-Kwan, Yu-Chiang Frank Wang +3
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
paper · doi:10.48550/arxiv.2512.22238
Abstract
Large-scale vision-language models (VLMs) have recently achieved remarkable multimodal understanding, but their massive size makes them impractical for deployment on mobile or edge devices. This raises the need for compact yet capable VLMs that can efficiently learn from powerful large teachers. However, distilling knowledge from a large teacher to a small student remains challenging due to their large size gap: the student often fails to reproduce the teacher's complex, high-dimensional representations, leading to unstable learning and degraded performance. To address this, we propose Masters (Masking Teacher and Reinforcing Student), a mask-progressive reinforcement learning (RL) distillation framework. Masters first masks non-dominant weights of the teacher to reduce unnecessary complexity, then progressively restores the teacher by gradually increasing its capacity during training. This strategy allows the student to learn richer representations from the teacher in a smooth and stable manner. To further refine knowledge transfer, Masters integrates an offline RL stage with two complementary rewards: an accuracy reward that measures the correctness of the generated responses, and a distillation reward that quantifies the ease of transferring responses from teacher to student. Unlike online think-answer RL paradigms that are computationally expensive and generate lengthy responses, our offline RL leverages pre-generated responses from masked teachers. These provide rich yet efficient guidance, enabling students to achieve strong performance without requiring the think-answer process.
Citations
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Kwai Keye-VL 1.5 Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- SmolVLM: Redefining small and efficient multimodal models
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition
- Qwen2.5-VL Technical Report
- Vision-Language Models for Edge Networks: A Comprehensive Survey
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Visual Large Language Models for Generalized and Specialized Applications
- MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
- Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation
- FastVLM: Efficient Vision Encoding for Vision Language Models
- OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
- Reasoning Beyond Words ? Exploring framework for hidden state reasoning
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
- GPT-4o System Card
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
- Phantom of Latent for Large Language and Vision Models
- NVLM: Open Frontier-Class Multimodal LLMs
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
- FuseChat: Knowledge Fusion of Chat Models
- Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
- LLaVA-OneVision: Easy Visual Task Transfer
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- LLAVADI: What Matters For Multimodal Large Language Models Distillation
- Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
- AMD: Automatic Multi-step Distillation of Large-scale Vision Models
- Dual-Space Knowledge Distillation for Large Language Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- TroL: Traversal of Layers for Large Language and Vision Models
- WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning
- M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
- MAmmoTH2: Scaling Instructions from the Web
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- MoVA: Adapting Mixture of Vision Experts to Multimodal Context
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Benchmarking Object Detectors with COCO: A New Path Forward
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
- VL-Mamba: Exploring State Space Models for Multimodal Learning
- MoAI: Mixture of All Intelligence for Large Language and Vision Models
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- TinyLLaVA: A Framework of Small-scale Large Multimodal Models
- CoLLaVO: Crayon Large Language and Vision mOdel
- Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
- MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
- G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Towards the Law of Capacity Gap in Distilling Language Models
- Open-Set Image Tagging with Multi-Grained Text Supervision
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- MMBench: Is Your Multi-modal Model an All-around Player?
- LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- CrossKD: Cross-Head Knowledge Distillation for Object Detection
- Scaling Open-Vocabulary Object Detection
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Lifting the Curse of Capacity Gap in Distilling Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Micrograph segmentations for DDEVD
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- Structured Pruning for Deep Convolutional Neural Networks: A Survey
- UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
- Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
- MapQA: A Dataset for Question Answering on Choropleth Maps
- EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
- Scaling Instruction-Finetuned Language Models
- Sparse Teachers Can Be Dense with Knowledge
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
- Panoptic Scene Graph Generation
- Visual Spatial Reasoning
- Masking Adversarial Damage: Finding Adversarial Saliency for Robust and Sparse Network
- PaLM: Scaling Language Modeling with Pathways
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- It's All in the Head: Representation Knowledge Distillation through Classifier Sharing
- Masked-attention Mask Transformer for Universal Image Segmentation
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System
- Scaling Vision with Sparse Mixture of Experts
- BERT Learns to Teach: Knowledge Distillation with Meta Learning
- Return-based Scaling: Yet Another Normalisation Trick for Deep RL
- InfographicVQA
- Distilling Knowledge via Knowledge Review
- Densely Guided Knowledge Distillation using Multiple Teacher Assistants
- DocVQA: A Dataset for VQA on Document Images
- Scaling Laws for Neural Language Models
- Contrastive Representation Distillation
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Patient Knowledge Distillation for BERT Model Compression
- Similarity-Preserving Knowledge Distillation
- OK-VQA: A Visual Question Answering Benchmark Requiring External\n Knowledge
- EigenDamage: Structured Pruning in the Kronecker-Factored Eigenbasis
- Relational Knowledge Distillation
- A Comprehensive Overhaul of Feature Distillation
- Improved Knowledge Distillation via Teacher Assistant
- TallyQA: Answering Complex Counting Questions
- DVQA: Understanding Data Visualizations via Question Answering
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Channel Pruning for Accelerating Very Deep Neural Networks
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- A Diagram Is Worth A Dozen Images
- Learning both Weights and Connections for Efficient Neural Networks
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- Learning Factored Representations in a Deep Mixture of Experts
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Related