LoRA Learns Less and Forgets Less
2024/05/15 by Dan Biderman, Biderman, Dan, Jacob Portes +21 · 3 voices · 116 citations
Engineering · Psychology · #Computer science #Economics #Psychology #Robotics and Automated Systems
paper · pdf · doi:10.48550/arxiv.2405.09673
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/05/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Low-Rank Adaptation (LoRA) is a widely-used parameter-efficient finetuning method for large language models. LoRA saves memory by training only low rank perturbations to selected weight matrices. In this work, we compare the performance of LoRA and full finetuning on two target domains, programming and mathematics. We consider both the instruction finetuning (approximately 100K prompt-response pairs) and continued pretraining (20B unstructured tokens) data regimes. Our results show that, in the standard low-rank settings, LoRA substantially underperforms full finetuning. Nevertheless, LoRA better maintains the base model's performance on tasks outside the target domain. We show that LoRA mitigates forgetting more than common regularization techniques such as weight decay and dropout; it also helps maintain more diverse generations. Finally, we show that full finetuning learns perturbations with a rank that is 10-100X greater than typical LoRA configurations, possibly explaining some of the reported gaps. We conclude by proposing best practices for finetuning with LoRA.
Cited by
- How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
- On the Convergence of Stochastic Low-Rank Adaptation
- PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer
- Scaling Point-in-Time Language Models
- Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
- Time Series Foundation Models for Process Model Forecasting
- The Appeal and Reality of Recycling LoRAs with Adaptive Merging
- seqLens: Optimizing Language Models for Genomic Predictions
- Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning
- The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning
- IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- TokenMem: Faithful Knowledge Injection for Frozen LLMs
- StAR: Segment Anything Reasoner
- On the Convergence Rate of LoRA Gradient Descent
- Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- How (Mis)calibrated is Your Federated CLIP and What To Do About It?
- Exploring the Rashomon Set for Concept-Based Models
- Parameter Importance-Driven Continual Learning for Foundation Models
- On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
- A mathematical theory of balancing relational generalization and memorization
- LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language
- Random Initialization of Gated Sparse Adapters
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Conditions for Catastrophic Forgetting in Multilingual Translation
- Continual Learning via Sparse Memory Finetuning
- OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
- Scaling Language-Centric Omnimodal Representation Learning
- CoLoR-GAN: Continual Few-Shot Learning with Low-Rank Adaptation in Generative Adversarial Networks
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- CTR-LoRA: Curvature-Aware and Trust-Region Guided Low-Rank Adaptation for Large Language Models
- How to Teach Large Multimodal Models New Skills
- Maximum In-Support Return Modeling for Dynamic Recommendation with Language Model Prior
- Revisiting Mixout: An Overlooked Path to Robust Finetuning
- Teamwork: Collaborative Diffusion with Low-rank Coordination and Adaptation
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- Optimizing Fine-Tuning through Advanced Initialization Strategies for Low-Rank Adaptation
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- Effective Quantization of Muon Optimizer States
- MolSpectLLM: A Molecular Foundation Model Bridging Spectroscopy, Molecule Elucidation, and 3D Structure Generation
- Unsupervised Defect Detection for Surgical Instruments
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- Latent Traits and Cross-Task Transfer: Deconstructing Dataset Interactions in LLM Fine-tuning
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- Continually Adding New Languages to Multilingual Language Models
- Singular Value Few-shot Adaptation of Vision-Language Models
- Domain Adaptation of LLMs for Process Data
- Incident Analysis for AI Agents
- LoRAtorio: An intrinsic approach to LoRA Skill Composition
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- Bayesian BiLO: Bilevel Local Operator Learning for Efficient Uncertainty Quantification of Bayesian PDE Inverse Problems with Low-Rank Adaptation
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
- Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving
- The Primacy of Magnitude in Low-Rank Adaptation
- T-LoRA: Single Image Diffusion Model Customization Without Overfitting
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimization
- Can Large Language Models Automate the Refinement of Cellular Network Specifications?
- Continual Gradient Low-Rank Projection Fine-Tuning for LLMs
- Beyond Low-Rank Tuning: Model Prior-Guided Rank Allocation for Effective Transfer in Low-Data and Large-Gap Regimes
- Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
- Pay Attention to Small Weights
- ReCode: Updating Code API Knowledge with Reinforcement Learning
- PrivacyXray: Detecting Privacy Breaches in LLMs through Semantic Consistency and Probability Certainty
- Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
- Improving LoRA with Variational Learning
- LARGO: Low-Rank Regulated Gradient Projection for Robust Parameter Efficient Fine-Tuning
- SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
- Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
- MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
- LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
- Continual Learning in Vision-Language Models via Aligned Model Merging
- MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
- Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution
- Weight Spectra Induced Efficient Model Adaptation
- LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear Mapping
- Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting
- Context-Free Synthetic Data Mitigates Forgetting
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation
- Adaptive parameter-efficient fine-tuning via Hessian-informed subset selection
- PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- Parameter Efficient Continual Learning with Dynamic Low-Rank Adaptation
- Cross-Benchmark Generalization in Long-Horizon Agents
- Achieving Scalable Robot Autonomy via neurosymbolic planning using lightweight local LLM
- (How) Learning Rates Regulate Catastrophic Overtraining
- Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
- Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
- ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
- Memorization and Knowledge Injection in Gated LLMs
- Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
- Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation
- KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- AROMA: Autonomous Rank-one Matrix Adaptation
- Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
- CSPLADE: Learned Sparse Retrieval with Causal Language Models
- GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions
Discussions
Related