Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
2025/08/06 by Yuyang Liu, Liu, Yuyang, Qiuhe Hong +14 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2508.04227
arxiv created 2026/07/30 · arxiv updated 2026/07/31
Abstract
Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot generalization. However, enabling them to learn continually from non-stationary data remains a major challenge, as their cross-modal alignment and generalization capabilities are particularly vulnerable to catastrophic forgetting. Unlike traditional unimodal continual learning (CL), VLMs face unique challenges such as cross-modal feature drift, parameter interference due to shared architectures, and zero-shot capability erosion. Furthermore, generative MLLMs exhibit a unique "alignment tax," where catastrophic forgetting manifests not merely as factual amnesia, but as a systemic collapse of deep Chain-of-Thought (CoT) reasoning. This survey presents the first comprehensive diagnostic review bridging continual learning across predictive VLMs and generative MLLMs. We systematically deconstruct the aforementioned failure modes and propose a challenge-driven taxonomy comprising four core paradigms: (1) Multi-Modal Replay Strategies addressing explicit and implicit memory drift; (2) Cross-Modal Regularization enforcing topological and geometric alignment; (3) Parameter-Efficient Adaptation utilizing dynamic routing and subspace projections; and the emerging (4) Model Fusion and Decoupling paradigms. We critically analyze the evolution of evaluation protocols, highlighting the essential shift toward dual-track benchmarks (Domain vs. Ability CL). Finally, we chart a roadmap for future research, emphasizing compositional zero-shot learning, embodied AI with sensor fusion, and autonomous agentic ecosystems. All resources are available at: https://github.com/YuyangSunshine/Awesome-Continual-learning-of-Vision-Language-Models
Citations
- Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- Prompt-Based Continual Compositional Zero-Shot Learning
- Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models
- Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning
- Closing the Modality Gap for Mixed Modality Search
- Overcoming catastrophic forgetting in neural networks
- Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- LLaVA-c: Continual Improved Visual Instruction Tuning
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
- Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
- SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
- Language Guided Concept Bottleneck Models for Interpretable Continual Learning
- HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
- Federated Continual Instruction Tuning
- Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
- LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
- Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
- CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
- ILIAS: Instance-Level Image retrieval At Scale
- Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
- The Double-Ellipsoid Geometry of CLIP
- Low-rank Prompt Interaction for Continual Vision-Language Retrieval
- Is Parameter Collision Hindering Continual Learning in LLMs?
- ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy
- ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
- Recent Advances of Multimodal Continual Learning: A Comprehensive Survey
- LW2G: Learning Whether to Grow for Prompt-based Continual Learning
- Anytime Continual Learning for Open Vocabulary Classification
- A Practitioner's Guide to Continual Multimodal Pretraining
- Boosting Open-Domain Continual Learning via Leveraging Intra-domain Category-aware Prototype
- Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
- CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
- Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion
- Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
- Advancing Cross-domain Discriminability in Continual Learning of Vision-Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- LoRA Learns Less and Forgets Less
- Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
- Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
- CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models
- Generative Multi-modal Models are Good Class-Incremental Learners
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
- CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning
- Select and Distill: Selective Dual-Teacher Knowledge Transfer for Continual Learning on Vision-Language Models
- CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
- PromptMM: Multi-Modal Knowledge Distillation for Recommendation with Prompt-Tuning
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
- Weighted Ensemble Models Are Strong Continual Learners
- Continual Instruction Tuning for Large Multimodal Models
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Class Incremental Learning with Pre-trained Vision-Language Models
- TiC-CLIP: Continual Training of CLIP Models
- Orthogonal Subspace Learning for Language Model Continual Learning
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Overcoming Generic Knowledge Loss with Selective Parameter Update
- ALIP: Adaptive Language-Image Pre-training with Synthetic Caption
- CTP: Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- AttriCLIP: A Non-Incremental Learner for Incremental Knowledge Learning
- Continual Multimodal Knowledge Graph Construction
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- Vision-Language Models for Vision Tasks: A Survey
- Vision-Language Models for Vision Tasks: A Survey
- Sigmoid Loss for Language Image Pre-Training
- A Unified Continual Learning Framework with General Parameter-Efficient Tuning
- Revisiting Class-Incremental Learning with Pre-Trained Models: Generalizability and Adaptivity are All You Need
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without Interference
- SuS-X: Training-Free Name-Only Transfer of Vision-Language Models
- Generative Negative Text Replay for Continual Vision-Language Pretraining
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental Learning
- Don't Stop Learning: Towards Continual Learning for the CLIP Model
- CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
- A Continual Deepfake Detection Benchmark: Dataset, Methods, and Essentials
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
- A Survey of Vision-Language Pre-Trained Models
- Rebalancing Batch Normalization for Exemplar-based Class-Incremental Learning
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- The CLEAR Benchmark: Continual LEArning on Real-World Imagery
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Learning to Prompt for Continual Learning
- FLAVA: A Foundational Language And Vision Alignment Model
- Grounded Language-Image Pre-training
- Towards a Unified View of Parameter-Efficient Transfer Learning
- Robust fine-tuning of zero-shot models
- Finetuned Language Models Are Zero-Shot Learners
- RECALL: Replay-based Continual Learning in Semantic Segmentation
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- LoRA: Low-Rank Adaptation of Large Language Models
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- Continual learning in cross-modal retrieval
- M6: A Chinese Multimodal Pretrainer
- Learning Transferable Visual Models From Natural Language Supervision
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize\n Long-Tail Visual Concepts
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- Incremental Few-Shot Object Detection
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Parameter-Efficient Transfer Learning for NLP
- Moment Matching for Multi-Source Domain Adaptation
- Representation Learning with Contrastive Predictive Coding
- Audio-Visual Event Localization in Unconstrained Videos
- Memory Aware Synapses: Learning what (not) to forget
- Lifelong Learning with Dynamically Expandable Networks
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Gradient Episodic Memory for Continual Learning
- CORe50: a New Dataset and Benchmark for Continuous Object Recognition
- Continual Learning with Deep Generative Replay
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- iCaRL: Incremental Classifier and Representation Learning
- Learning without Forgetting
- Progressive Neural Networks
- VQA: Visual Question Answering
- Microsoft COCO: Common Objects in Context
- Adaptive Mixtures of Local Experts
- Take Only What You Need: Rank Minimization as an Implicit Forgetting Regularizer in Continual Learning
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
Cited by
Related