vix.ing · top · new · best · stats · spec

Distill, Adapt, Distill: Training Small, In-Domain Models for Neural\n Machine Translation

2020/03/05 by Mitchell A. Gordon, Kevin Duh, Gordon, Mitchell A. +1
Computer Science · Mathematics · Psychology · #Adaptation (eye) #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Distillation #Domain (mathematical analysis) #Domain adaptation #Domain knowledge #FOS: Computer and information sciences #Language model #Machine learning #Machine translation #Mathematics #Natural Language Processing Techniques #Natural language processing #Psychology #Scale (ratio) #Text Readability and Simplification #Topic Modeling #Training set #Translation (biology)

paper · pdf · doi:10.48550/arxiv.2003.02877

openalex publication_date 2020/03/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We explore best practices for training small, memory efficient machine\ntranslation models with sequence-level knowledge distillation in the domain\nadaptation setting. While both domain adaptation and knowledge distillation are\nwidely-used, their interaction remains little understood. Our large-scale\nempirical results in machine translation (on three language pairs with three\ndomains each) suggest distilling twice for best performance: once using\ngeneral-domain data and again using in-domain data with an adapted teacher.\n

Related