Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
2023/11/06 by Le Yu, Yu, Le, Bowen Yu +7 · 1 voice · 204 citations
Computer Science · #Algorithm #Artificial intelligence #Computer science #Encoder #Language model #Merge (version control) #Natural Language Processing Techniques #Parallel computing #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2311.03099
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In this paper, we unveil that Language Models (LMs) can acquire new capabilities by assimilating parameters from homologous models without retraining or GPUs. We first introduce DARE to set most delta parameters (i.e., the disparity between fine-tuned and pre-trained parameters) to zeros without affecting the abilities of Supervised Fine-Tuning (SFT) LMs, which randomly Drops delta parameters with a ratio p And REscales the remaining ones by 1 / (1 - p) to approximate the original embeddings. Then, we use DARE as a versatile plug-in to sparsify delta parameters of multiple SFT homologous models for mitigating parameter interference and merge them into a single model by parameter fusing. We experiment with encoder- and decoder-based LMs, showing that: (1) SFT delta parameter value ranges are typically small (within 0.002) with extreme redundancy, and DARE can effortlessly eliminate 90% or even 99% of them; (2) DARE can merge multiple task-specific LMs into one LM with diverse capabilities. Notably, this phenomenon is more pronounced in large-scale LMs, where the merged LM reveals the potential to surpass the performance of any source LM, providing a new discovery. We also utilize DARE to create a merged LM that ranks first among models with 7 billion parameters on the Open LLM Leaderboard.
Cited by
- Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training
- Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
- BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors
- CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging
- Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach
- When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
- Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Making Open-Source Text LLM Watermarks Durable Against Merging
- The Universal Weight Subspace Hypothesis
- TRINITY: An Evolved LLM Coordinator
- Competition and Attraction Improve Model Fusion
- Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
- Model Merging via Multi-Teacher Knowledge Distillation
- FaithLens: Detecting and Explaining Faithfulness Hallucination
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- Bridging Training and Merging Through Momentum-Aware Optimization
- Null-LoRA: Low-Rank Adaptation on Null Space
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
- AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
- Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- EvoEdit: Lifelong Free-Text Knowledge Editing through Latent Perturbation Augmentation and Knowledge-driven Parameter Fusion
- Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
- Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
- A Systematic Study of In-the-Wild Model Merging for Large Language Models
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Defending Unauthorized Model Merging via Dual-Stage Weight Protection
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
- Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
- WeaveRec: An LLM-Based Cross-Domain Sequential Recommendation Framework with Model Merging
- Stemma: Induced Decision Regions Reveal LLM Provenance
- Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- MIN-Merging: Merge the Important Neurons for Model Merging
- World Simulation with Video Foundation Models for Physical AI
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- Model Merging with Functional Dual Anchors
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
- Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
- Directional Reasoning Injection for Fine-Tuning MLLMs
- Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
- Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- Towards Reversible Model Merging For Low-rank Weights
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
- Weight Weaving: Parameter Pooling for Data-Free Model Merging
- Exploring and Leveraging Class Vectors for Classifier Editing
- Revisiting Model Interpolation for Efficient Reasoning
- Don't Throw Away Your Pretrained Model
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- NurseLLM: The First Specialized Language Model for Nursing
- A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
- BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
- FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- Model Merging Scaling Laws in Large Language Models
- The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
- Effect of Model Merging in Domain-Specific Ad-hoc Retrieval
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Transparency, Promoting Fidelity
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- Personality Vector: Modulating Personality of Large Language Models by Model Merging
- Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
- AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
- A Pipeline to Assess Merging Methods via Behavior and Internals
- Introducing LongCat-Flash-Thinking: A Technical Report
- Accurate and Efficient Low-Rank Model Merging in Core Space
- SEQR: Secure and Efficient QR-based LoRA Routing
- TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
- Variational Task Vector Composition
- HAM: Hierarchical Adapter Merging for Scalable Continual Learning
- Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- Continually Adding New Languages to Multilingual Language Models
- Merge-of-Thought Distillation
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- Surrogate Benchmarks for Model Merging Optimization
- DivMerge: A divergence-based model merging method for multi-tasking
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- On Task Vectors and Gradients
- Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing
- Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models
- Rethinking Layer-wise Model Merging through Chain of Merges
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- Efficient Multi-Source Knowledge Transfer by Model Merging
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
- Grounding Multilingual Multimodal LLMs With Cultural Knowledge
- Low-Rank Expert Merging for Multi-Source Domain Adaptation in Person Re-Identification
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- Tensorized Clustered LoRA Merging for Multi-Task Interference
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- Efficient Compositional Multi-tasking for On-device Large Language Models
- MeMo: Memory as a Model
- Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning
- HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
- RegCL: Continual Adaptation of Segment Anything Model via Model Merging
- Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
- Selecting and Merging: Towards Adaptable and Scalable Named Entity Recognition with Large Language Models
- DuET: Dual Incremental Object Detection via Exemplar-Free Task Arithmetic
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging
- Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
- From Memorization to Parameter Interference: How Overtraining Experts Harms Model Merging
- MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task Adaptation
- Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
- Position: Pause Recycling LoRAs and Prioritize Mechanisms to Uncover Limits and Effectiveness
- MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
- A correlation-permutation approach for speech-music encoders model merging
- ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- STELLA: A Multimodal LLM for Protein Functional Annotation via Unified Sequence-Structure Encoding
- Bohdi: Heterogeneous LLM Fusion with Automatic Data Exploration
- FroM: Frobenius Norm-Based Data-Free Adaptive Model Merging
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?
- FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA
- Assembly of Experts: Linear-time construction of the Chimera LLM variants with emergent and adaptable behaviors
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
- Navigating the Accuracy-Size Trade-Off with Flexible Model Merging
- Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking
- Permissioned LLMs: Enforcing Access Control in Large Language Models
- LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
- Why Do More Experts Fail? A Theoretical Analysis of Model Merging
- SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
- The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
- Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
- Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
- ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation
- Composable Cross-prompt Essay Scoring by Merging Models
- Knowledge Grafting of Large Language Models
- The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
- AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model
- Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
- Training-Free Reasoning and Reflection in MLLMs
- Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
- NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
- Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
- DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
- Merge to Mix: Mixing Datasets via Model Merging
- Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
- Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
- Activation-Guided Consensus Merging for Large Language Models
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
- SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation
- Scalable Strategies for Continual Learning with Replay
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
- Model Merging in Pre-training of Large Language Models
- LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
- Mergenetic: a Simple Evolutionary Model Merging Library
- MergeBench: A Benchmark for Merging Domain-Specialized LLMs
- Dynamic Base model Shift for Delta Compression
- RanDeS: Randomized Delta Superposition for Multi-Model Compression
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
- HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- Scalable Token-Level Hallucination Detection in Large Language Models
- Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
- Dynamic Fisher-weighted Model Merging via Bayesian Optimization
- χ0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies
- Per-parameter Task Arithmetic for Unlearning in Large Language Models
- Multi-task Code LLMs: Data Mix or Model Merge?
- SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
- A Model Merging Approach for Continual MLLM Unlearning
- Suppression Sticks, Locality Is Fragile: A Closed-Loop Target-and-Control Audit of Task-Vector Negation in VLA Policies
- EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
- A Unified Model for Cross-Domain Clone Detection via Model Merging
- MergeSE: Post-Hoc Model Merging for Software Engineering Tasks Without Retraining
- Exact Unlearning of Finetuning Data via Model Merging at Scale
- MASS: MoErging through Adaptive Subspace Selection
- EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
- Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
- Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-Tuning
- ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
- Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
- Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
- LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
- Defending Deep Neural Networks against Backdoor Attacks via Module Switching
- SEA-LION: Southeast Asian Languages in One Network
Discussions
Related