2020/04/15 by Chien-Sheng Wu, Steven Hoi, Steven C. H. Hoi +6 · 7 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech and dialogue systems #Text Readability and Simplification #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2004.06871
EMNLP 2020 camera-ready
openalex publication_date 2020/04/15 · arxiv created 2020/10/01 · arxiv updated 2020/10/02 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
The underlying difference of linguistic patterns between general text and task-oriented dialogue makes existing pre-trained language models less useful in practice. In this work, we unify nine human-human and multi-turn task-oriented dialogue datasets for language modeling. To better model dialogue behavior during pre-training, we incorporate user and system tokens into the masked language modeling. We propose a contrastive objective function to simulate the response selection task. Our pre-trained task-oriented dialogue BERT (TOD-BERT) outperforms strong baselines like BERT on four downstream task-oriented dialogue applications, including intention recognition, dialogue state tracking, dialogue act prediction, and response selection. We also show that TOD-BERT has a stronger few-shot ability that can mitigate the data scarcity problem for task-oriented dialogue.