TIES-Merging: Resolving Interference When Merging Models
2023/06/02 by Prateek Yadav, Yadav, Prateek, Derek Tam +7 · 2 voices · 250 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Topic Modeling #cs.AI #cs.CL #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2306.01708
arxiv published 2023/06/02 · arxiv updated 2023/10/27
Abstract
Transfer learning - i.e., further fine-tuning a pre-trained model on a downstream task - can confer significant advantages, including improved downstream performance, faster convergence, and better sample efficiency. These advantages have led to a proliferation of task-specific fine-tuned models, which typically can only perform a single task and do not benefit from one another. Recently, model merging techniques have emerged as a solution to combine multiple task-specific models into a single multitask model without performing additional training. However, existing merging methods often ignore the interference between parameters of different models, resulting in large performance drops when merging multiple models. In this paper, we demonstrate that prior merging techniques inadvertently lose valuable information due to two major sources of interference: (a) interference due to redundant parameter values and (b) disagreement on the sign of a given parameter's values across models. To address this, we propose our method, TRIM, ELECT SIGN & MERGE (TIES-Merging), which introduces three novel steps when merging models: (1) resetting parameters that only changed a small amount during fine-tuning, (2) resolving sign conflicts, and (3) merging only the parameters that are in alignment with the final agreed-upon sign. We find that TIES-Merging outperforms several existing methods in diverse settings covering a range of modalities, domains, number of tasks, model sizes, architectures, and fine-tuning settings. We further analyze the impact of different types of interference on model parameters, and highlight the importance of resolving sign interference. Our code is available at https://github.com/prateeky2806/ties-merging
Cited by
- Rethinking Expert Training for Model Merging with Prompt Learning
- Mechanistic Evidence for Preserved-but-Misaligned Representations in Non-IID FedAvg
- Merge before Forget: A Single LoRA Continual Learning via Continual Merging
- Decomposing Task Vectors for Refined Model Editing
- GLUE: Gradient-free Learning to Unify Experts
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
- Model Merging via Multi-Teacher Knowledge Distillation
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- Bridging Training and Merging Through Momentum-Aware Optimization
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
- AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
- System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
- Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptation
- An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
- Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
- From Coefficients to Directions: Rethinking Model Merging with Directional Alignment
- Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
- A Systematic Study of In-the-Wild Model Merging for Large Language Models
- Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- Escaping Optimization Stagnation: Taking Steps Beyond Task Arithmetic via Difference Vectors
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
- RETROFIT: Continual Learning with Controlled Forgetting for Binary Security Detection and Analysis
- Defending Unauthorized Model Merging via Dual-Stage Weight Protection
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- Steering Language Models with Weight Arithmetic
- Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
- Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
- Parameterized Prompt for Incremental Object Detection
- T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
- WeaveRec: An LLM-Based Cross-Domain Sequential Recommendation Framework with Model Merging
- Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- MIN-Merging: Merge the Important Neurons for Model Merging
- World Simulation with Video Foundation Models for Physical AI
- Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
- Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
- Model Merging with Functional Dual Anchors
- Mapping Post-Training Forgetting in Language Models at Scale
- Hierarchical Federated Unlearning for Large Language Models
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
- Backdoor Unlearning by Linear Task Decomposition
- Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
- Towards Reversible Model Merging For Low-rank Weights
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
- Weight Weaving: Parameter Pooling for Data-Free Model Merging
- Exploring and Leveraging Class Vectors for Classifier Editing
- On-device System of Compositional Multi-tasking in Large Language Models
- REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration
- Decoupled DiLoCo for Resilient Distributed Pre-training
- Don't Throw Away Your Pretrained Model
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL
- Do We Really Need Permutations? Impact of Model Width on Linear Mode Connectivity
- Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
- How does the optimizer implicitly bias the model merging loss landscape?
- FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- Model Merging Scaling Laws in Large Language Models
- Merge Now, Regret Later: The Hidden Cost of Model Merging is Adversarial Transferability
- Toward a Holistic Approach to Continual Model Merging
- Temporal Generalization: A Reality Check
- The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Transparency, Promoting Fidelity
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
- Transporting Task Vectors across Different Architectures without Training
- AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
- Introducing LongCat-Flash-Thinking: A Technical Report
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
- Accurate and Efficient Low-Rank Model Merging in Core Space
- SEQR: Secure and Efficient QR-based LoRA Routing
- Variational Task Vector Composition
- HAM: Hierarchical Adapter Merging for Scalable Continual Learning
- Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
- Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
- Routing Distilled Knowledge via Mixture of LoRA Experts for Large Language Model based Bundle Generation
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- Continually Adding New Languages to Multilingual Language Models
- VARCO-VISION-2.0 Technical Report
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Surrogate Benchmarks for Model Merging Optimization
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- On Task Vectors and Gradients
- Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- Rethinking Layer-wise Model Merging through Chain of Merges
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning
- Efficient Multi-Source Knowledge Transfer by Model Merging
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- Cost-Aware Contrastive Routing for LLMs
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Low-Rank Expert Merging for Multi-Source Domain Adaptation in Person Re-Identification
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
- Neuro-MoBRE: Exploring Multi-subject Multi-task Intracranial Decoding via Explicit Heterogeneity Resolving
- Tensorized Clustered LoRA Merging for Multi-Task Interference
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- Test-Time Model Adaptation for Quantized Neural Networks
- RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Forgetting of task-specific knowledge in model merging-based continual learning
- Efficient Compositional Multi-tasking for On-device Large Language Models
- ReCatcher: Towards LLMs Regression Testing for Code Generation
- modelDNA: Calibrated Lineage Verification and Merge Decomposition from Sampled Weight Fingerprints
- Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
- Continual Speaker Identity Unlearning with Minimal Interference
- MeMo: Memory as a Model
- Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning
- HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
- Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
- RegCL: Continual Adaptation of Segment Anything Model via Model Merging
- Harmonizing and Merging Source Models for CLIP-based Domain Generalization
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- FlexOlmo: Open Language Models for Flexible Data Use
- Intrinsic Training Signals for Federated Learning Aggregation
- Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
- Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust Planning
- When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
- Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
- Eigenvoice Synthesis based on Model Editing for Speaker Generation
- CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
- Reducing Variability of Multiple Instance Learning Methods for Digital Pathology
- DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive Segmentation
- DuET: Dual Incremental Object Detection via Exemplar-Free Task Arithmetic
- Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- Command-V: Pasting LLM Behaviors via Activation Profiles
- SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging
- Subspace-Boosted Model Merging
- From Memorization to Parameter Interference: How Overtraining Experts Harms Model Merging
- MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task Adaptation
- Position: Pause Recycling LoRAs and Prioritize Mechanisms to Uncover Limits and Effectiveness
- Training-free LLM Merging for Multi-task Learning
- Generative Representational Learning of Foundation Models for Recommendation
- A correlation-permutation approach for speech-music encoders model merging
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
- StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
- Bohdi: Heterogeneous LLM Fusion with Automatic Data Exploration
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
- FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA
- Assembly of Experts: Linear-time construction of the Chimera LLM variants with emergent and adaptable behaviors
- Continual Learning in Vision-Language Models via Aligned Model Merging
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
- Navigating the Accuracy-Size Trade-Off with Flexible Model Merging
- Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking
- Permissioned LLMs: Enforcing Access Control in Large Language Models
- Learning Composable Chains-of-Thought
- Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation
- Update Your Transformer to the Latest Release: Re-Basin of Task Vectors
- LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
- Why Do More Experts Fail? A Theoretical Analysis of Model Merging
- FCOS: A Two-Stage Recoverable Model Pruning Framework for Automatic Modulation Recognition
- Multi-objective Large Language Model Alignment with Hierarchical Experts
- SeMe: Training-Free Language Model Merging via Semantic Alignment
- Bielik 11B v2 Technical Report
- The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
- Composable Cross-prompt Essay Scoring by Merging Models
- Knowledge Grafting of Large Language Models
- Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
- Training-Free Reasoning and Reflection in MLLMs
- CodeMerge: Codebook-Guided Model Merging for Robust Test-Time Adaptation in Autonomous Driving
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
- Merge to Mix: Mixing Datasets via Model Merging
- Model Merging is Secretly Certifiable: Non-Vacuous Generalisation Bounds for Low-Shot Learning
- Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
- Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
- Activation-Guided Consensus Merging for Large Language Models
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
- Text Generation Beyond Discrete Token Sampling
- Distilling a speech and music encoder with task arithmetic
- SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation
- Scalable Strategies for Continual Learning with Replay
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
- Model Merging in Pre-training of Large Language Models
- Mergenetic: a Simple Evolutionary Model Merging Library
- MergeBench: A Benchmark for Merging Domain-Specialized LLMs
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
- Dynamic Base model Shift for Delta Compression
- RanDeS: Randomized Delta Superposition for Multi-Model Compression
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- Customizing a Large Language Model for VHDL Design of High-Performance Microprocessors
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model Merging
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment
- AI-Model Network: Concept, Current State and Future
- Bus-Conditioned Zero-Shot Trajectory Generation via Task Arithmetic
- Shared LoRA Subspaces for almost Strict Continual Learning
- MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
- DanceOPD: On-Policy Generative Field Distillation
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
- Scalable Token-Level Hallucination Detection in Large Language Models
- Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
- Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
- χ0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies
- Per-parameter Task Arithmetic for Unlearning in Large Language Models
- Multi-task Code LLMs: Data Mix or Model Merge?
- SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
- A Model Merging Approach for Continual MLLM Unlearning
- Suppression Sticks, Locality Is Fragile: A Closed-Loop Target-and-Control Audit of Task-Vector Negation in VLA Policies
- EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
- When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
- A Unified Model for Cross-Domain Clone Detection via Model Merging
- MergeSE: Post-Hoc Model Merging for Software Engineering Tasks Without Retraining
- Exact Unlearning of Finetuning Data via Model Merging at Scale
- MASS: MoErging through Adaptive Subspace Selection
- EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
- Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-Tuning
- Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
- Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
- How new data permeates LLM knowledge and how to dilute it
- LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
- Defending Deep Neural Networks against Backdoor Attacks via Module Switching
- SpectR: Dynamically Composing LM Experts with Spectral Routing
Discussions
Related