vix.ing · top · new · best · stats · spec

Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

2024/11/11 by Megh Thakkar, Thakkar, Megh, Quentin Fournier +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Digital Rights Management and Security #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2411.06824

openalex publication_date 2024/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models often experience a loss in their safety abilities in the process, making them capable of generating harmful content. As a solution, we introduce an efficient and effective merging-based alignment method called MergeAlign that interpolates the domain and alignment vectors, creating safer domain-specific models while preserving their utility. We apply MergeAlign on Llama3 variants that are experts in medicine and finance, obtaining substantial alignment improvements with minimal to no degradation on domain-specific benchmarks. We study the impact of model merging through model similarity metrics and contributions of individual models being merged. We hope our findings open new research avenues and inspire more efficient development of safe expert LLMs.

Cited by

Related