vix.ing · top · new · best · stats

Improving BERT Fine-Tuning via Self-Ensemble and Self-Distillation

2020/02/24 by Yige Xu, Xipeng Qiu, Xu, Yige +5 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2002.10345

7 pages, 6 figures

arxiv created 2020/02/24 · arxiv updated 2020/02/25

Abstract

Fine-tuning pre-trained language models like BERT has become an effective way in NLP and yields state-of-the-art results on many downstream tasks. Recent studies on adapting BERT to new tasks mainly focus on modifying the model structure, re-designing the pre-train tasks, and leveraging external data and knowledge. The fine-tuning strategy itself has yet to be fully explored. In this paper, we improve the fine-tuning of BERT with two effective mechanisms: self-ensemble and self-distillation. The experiments on text classification and natural language inference tasks show our proposed methods can significantly improve the adaption of BERT without any external data or knowledge.

Cited by

Related