vix.ing · top · new · best · stats

Multi-Stage Balanced Distillation: Addressing Long-Tail Challenges in Sequence-Level Knowledge Distillation

2024/06/19 by Yuhang Zhou, Zhou, Yuhang, Jing Zhu +13 · 2 citations
Chemistry · Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Biology #Chemistry #Chromatography #Computation and Language (cs.CL) #Computer science #Distillation #Engineering #Environmental science #FOS: Computer and information sciences #Industrial engineering #Machine Learning and Algorithms #Process Optimization and Integration #Process engineering #Reservoir Engineering and Simulation Methods #Sequence (biology) #Stage (stratigraphy)

paper · pdf · doi:10.48550/arxiv.2406.13114

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/06/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

Large language models (LLMs) have significantly advanced various natural language processing tasks, but deploying them remains computationally expensive. Knowledge distillation (KD) is a promising solution, enabling the transfer of capabilities from larger teacher LLMs to more compact student models. Particularly, sequence-level KD, which distills rationale-based reasoning processes instead of merely final outcomes, shows great potential in enhancing students' reasoning capabilities. However, current methods struggle with sequence level KD under long-tailed data distributions, adversely affecting generalization on sparsely represented domains. We introduce the Multi-Stage Balanced Distillation (BalDistill) framework, which iteratively balances training data within a fixed computational budget. By dynamically selecting representative head domain examples and synthesizing tail domain examples, BalDistill achieves state-of-the-art performance across diverse long-tailed datasets, enhancing both the efficiency and efficacy of the distilled models.

Cited by

Related