Multilingual E5 Text Embeddings: A Technical Report
2024/02/08 by Liang Wang, Nan Yang, Wang, Liang +9 · 82 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2402.05672
openalex publication_date 2024/02/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
This technical report presents the training methodology and evaluation results of the open-source multilingual E5 text embedding models, released in mid-2023. Three embedding models of different sizes (small / base / large) are provided, offering a balance between the inference efficiency and embedding quality. The training procedure adheres to the English E5 model recipe, involving contrastive pre-training on 1 billion multilingual text pairs, followed by fine-tuning on a combination of labeled datasets. Additionally, we introduce a new instruction-tuned embedding model, whose performance is on par with state-of-the-art, English-only models of similar sizes. Information regarding the model release can be found at https://github.com/microsoft/unilm/tree/master/e5 .
Cited by
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- Do LLM Debates Repeat Arguments Differently Across Languages?
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
- Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register
- TCDE: Topic-Centric Dual Expansion of Queries and Documents with Large Language Models for Information Retrieval
- Log Anomaly Detection with Large Language Models via Knowledge-Enriched Fusion
- Are Large Language Models Really Effective for Training-Free Cold-Start Recommendation?
- BAID: A Benchmark for Bias Assessment of AI Detectors
- CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale
- Unveiling Affective Polarization Trends in Parliamentary Proceedings
- Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
- SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
- From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
- uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
- What Drives Cross-lingual Ranking? Retrieval Approaches with Multilingual Language Models
- Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
- PolicyBot - Reliable Question Answering over Policy Documents
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- Evaluating Embedding Generalization: How LLMs, LoRA, and SLERP Shape Representational Geometry
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
- Query Generation Pipeline with Enhanced Answerability Assessment for Financial Information Retrieval
- EncouRAGe: Evaluating RAG Local, Fast, and Reliable
- Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
- GigaEmbeddings: Efficient Russian Language Embedding Model
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
- CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
- CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
- Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
- A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation
- Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- Exploring Instruction Data Quality for Explainable Image Quality Assessment
- Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
- A-VERT: Agnostic Verification with Embedding Ranking Targets
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
- Estimating the Empowerment of Language Model Agents
- Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
- The Medium Is Not the Message: Deconfounding Text Embeddings via Linear Concept Erasure
- An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?
- The hunt for the last relevant paper: blending the best of humans and AI
- ST-Raptor: LLM-Powered Semi-Structured Table Question Answering
- Chunk Knowledge Generation Model for Enhanced Information Retrieval: A Multi-task Learning Approach
- Evaluating Large Language Models for Cross-Lingual Retrieval
- Measuring Gender Bias in Job Title Matching for Grammatical Gender Languages
- Case-Based Decision-Theoretic Decoding with Quality Memories
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- THEME: Enhancing Thematic Investing with Semantic Stock Representations and Temporal Dynamics
- Modeling shopper interest broadness with entropy-driven dialogue policy in the context of arbitrarily large product catalogs
- MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages
- The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- Native Logical and Hierarchical Representations with Subspace Embeddings
- Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
- Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- MHSNet:An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language Model
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- Transforming Questions and Documents for Semantically Aligned Retrieval-Augmented Generation
- LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- Reliable Evaluation Protocol for Low-Precision Retrieval
- AIAP: A No-Code Workflow Builder for Non-Experts with Natural Language and Multi-Agent Collaboration
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- Zero-Shot Document Understanding using Pseudo Table of Contents-Guided Retrieval-Augmented Generation
- Real-time News Story Identification
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
Related