vix.ing · top · new · best · stats

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

2020/03/24 by Junjie Hu, Hu, Junjie, Sebastian Ruder +10 · 95 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2003.11080

In Proceedings of the 37th International Conference on Machine Learning (ICML). July 2020

openalex publication_date 2020/03/24 · arxiv created 2020/09/04 · arxiv updated 2020/09/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders XTREME benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.

Citations

Cited by

Related