vix.ing · top · new · best · stats

Better Fine-Tuning by Reducing Representational Collapse

2020/08/06 by Armen Aghajanyan, Akshat Shrivastava, Aghajanyan, Armen +9 · 9 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Topic Modeling #cs.CL #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2008.03156

arxiv created 2020/08/06 · openalex publication_date 2020/08/06 · arxiv updated 2020/08/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Although widely adopted, existing approaches for fine-tuning pre-trained language models have been shown to be unstable across hyper-parameter settings, motivating recent work on trust region methods. In this paper, we present a simplified and efficient method rooted in trust region theory that replaces previously used adversarial objectives with parametric noise (sampling from either a normal or uniform distribution), thereby discouraging representation change during fine-tuning when possible without hurting performance. We also introduce a new analysis to motivate the use of trust region methods more generally, by studying representational collapse; the degradation of generalizable representations from pre-trained models as they are fine-tuned for a specific end task. Extensive experiments show that our fine-tuning method matches or exceeds the performance of previous trust region methods on a range of understanding and generation tasks (including DailyMail/CNN, Gigaword, Reddit TIFU, and the GLUE benchmark), while also being much faster. We also show that it is less prone to representation collapse; the pre-trained models maintain more generalizable representations every time they are fine-tuned.

Citations

Cited by

Related