vix.ing · top · new · best · stats

Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models

2024/05/08 by Luke Merrick, Merrick, Luke, Danmei Xu +5 · 1 voice · 30 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.IR

paper · pdf · doi:10.48550/arxiv.2405.05374

openalex publication_date 2024/05/08 · arxiv published 2024/05/08 · arxiv updated 2024/05/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This report describes the training dataset creation and recipe behind the family of arctic-embed text embedding models (a set of five models ranging from 22 to 334 million parameters with weights open-sourced under an Apache-2 license). At the time of their release, each model achieved state-of-the-art retrieval accuracy for models of their size on the MTEB Retrieval leaderboard, with the largest model, arctic-embed-l outperforming closed source embedding models such as Cohere's embed-v3 and Open AI's text-embed-3-large. In addition to the details of our training recipe, we have provided several informative ablation studies, which we believe are the cause of our model performance.

Cited by

Discussions

Related