XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
2020/03/24 by Junjie Hu, Hu, Junjie, Sebastian Ruder +10 · 95 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2003.11080
In Proceedings of the 37th International Conference on Machine Learning (ICML). July 2020
openalex publication_date 2020/03/24 · arxiv created 2020/09/04 · arxiv updated 2020/09/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders XTREME benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.
Citations
Cited by
- Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
- BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
- How Language Directions Align with Token Geometry in Multilingual LLMs
- Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
- On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
- Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
- LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
- TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- LiRA: Linguistic Robust Anchoring for Cross-lingual Large Language Models
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- Large Language Models Hallucination: A Comprehensive Survey
- Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
- Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
- MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages
- Just Use XML: Revisiting Joint Translation and Label Projection
- Quantifying Language Disparities in Multilingual Large Language Models
- General Demographic Foundation Models for Enhancing Predictive Performance Across Diseases and Populations
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
- KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- chDzDT: Word-level morphology-aware language model for Algerian social media text
- Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answering
- GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
- When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
- Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
- Improving Generative Cross-lingual Aspect-Based Sentiment Analysis with Constrained Decoding
- Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
- Cross-Prompt Encoder for Low-Performing Languages
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- UWB at WASSA-2024 Shared Task 2: Cross-lingual Emotion Detection
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- Enhancing Hindi NER in Low Context: A Comparative study of Transformer-based models with vs. without Retrieval Augmentation
- From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
- Benchmarking LLM Privacy Recognition for Social Robot Decision Making
- Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?
- AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
- McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
- Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource Languages
- Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
- The Cognate Data Bottleneck in Language Phylogenetics
- Multilingual BERT Post-Pretraining Alignment
- MergeDistill: Merging Pre-trained Language Models using Distillation
- skLEP: A Slovak General Language Understanding Benchmark
- Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
- Unbiased Sentence Encoder For Large-Scale Multi-lingual Search Engines
- GLGE: A New General Language Generation Evaluation Benchmark
- From Zero to Hero: On the Limitations of Zero-Shot Cross-Lingual Transfer with Multilingual Transformers
- Mitigating Negative Interference in Multilingual Sequential Knowledge Editing through Null-Space Constraints
- A Culturally Rich Romanian NLP Dataset from 'Who Wants to Be a Millionaire?' Videos
- Robust Optimization for Multilingual Translation with Imbalanced Data
- Distilling Large Language Models into Tiny and Effective Students using pQRNN
- DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead
- The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
- Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
- On the Language-specificity of Multilingual BERT and the Impact of Fine-tuning
- On Multilingual Encoder Language Model Compression for Low-Resource Languages
- On Learning Universal Representations Across Languages
- Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
- ParsiNLU: A Suite of Language Understanding Challenges for Persian
- Phrase-level Active Learning for Neural Machine Translation
- MAPS: A Multilingual Benchmark for Agent Performance and Security
- Cross-Linguistic Transfer in Multilingual NLP: The Role of Language Families and Morphology
- JNLP at SemEval-2025 Task 11: Cross-Lingual Multi-Label Emotion Detection Using Generative Models
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- An Improved Framework for Scaling Party Positions from Texts with Transformer
- Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
- Cross-lingual Extended Named Entity Classification of Wikipedia Articles
- The Devil Is in the Word Alignment Details: On Translation-Based Cross-Lingual Transfer for Token Classification Tasks
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- MuRIL: Multilingual Representations for Indian Languages
- Automatic Construction of Evaluation Suites for Natural Language Generation Datasets
- Learning Multilingual Representation for Natural Language Understanding with Enhanced Cross-Lingual Supervision
- What Causes Knowledge Loss in Multilingual Language Models?
- Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation
- Problems and Countermeasures in Natural Language Processing Evaluation
- CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents
- Long-context Non-factoid Question Answering in Indic Languages
- Bias Beyond English: Evaluating Social Bias and Debiasing Methods in a Low-Resource Setting
- Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
Related