ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
2021/12/02 by Keshav Santhanam, Omar Khattab, Santhanam, Keshav +7 · 2 voices · 131 citations
Computer Science · #Information Retrieval and Search Behavior #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.2112.01488
NAACL 2022. Omar and Keshav contributed equally to this work
arxiv created 2022/07/10 · arxiv updated 2022/07/12
Abstract
Neural information retrieval (IR) has greatly advanced search and other knowledge-intensive language tasks. While many neural IR methods encode queries and documents into single-vector representations, late interaction models produce multi-vector representations at the granularity of each token and decompose relevance modeling into scalable token-level computations. This decomposition has been shown to make late interaction more effective, but it inflates the space footprint of these models by an order of magnitude. In this work, we introduce ColBERTv2, a retriever that couples an aggressive residual compression mechanism with a denoised supervision strategy to simultaneously improve the quality and space footprint of late interaction. We evaluate ColBERTv2 across a wide range of benchmarks, establishing state-of-the-art quality within and outside the training domain while reducing the space footprint of late interaction models by 6--10×.
Cited by
- Hallucination Detection and Evaluation of Large Language Model
- Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval
- JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
- brat: Aligned Multi-View Embeddings for Brain MRI Analysis
- PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering
- Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
- Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- Incorporating Token Importance in Multi-Vector Retrieval
- TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- QueStER: Query Specification for Generative keyword-based Retrieval
- An Efficient Proximity Graph-based Approach to Table Union Search
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy
- LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026
- LLM generation novelty through the lens of semantic similarity
- The novel vector database
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
- DSEBench: A Test Collection for Explainable Dataset Search with Examples
- BiMax: Bidirectional MaxSim Score for Document-Level Alignment
- Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report
- Probing Latent Knowledge Conflict for Faithful Retrieval-Augmented Generation
- Simple Projection Variants Improve ColBERT Performance
- Scalable In-context Ranking with Generative Models
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- Bridging Clinical Narratives and ACR Appropriateness Guidelines: A Multi-Agent RAG System for Medical Imaging Decisions
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- Less LLM, More Documents: Searching for Improved RAG
- Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
- DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
- G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge
- Investigating Multi-layer Representations for Dense Passage Retrieval
- QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance
- EmbeddingGemma: Powerful and Lightweight Text Representations
- GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
- Retrieval from Within: An Intrinsic Capability of Attention-Based Models
- Semantic Search for Information Retrieval
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- Multi-view-guided Passage Reranking with Large Language Models
- Identifying Origins of Place Names via Retrieval Augmented Generation
- KG-CQR: Leveraging Structured Relation Representations in Knowledge Graphs for Contextual Query Retrieval
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
- Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering
- CLAP: Coreference-Linked Augmentation for Passage Retrieval
- BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
- MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
- DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
- Code Clone Detection via an AlphaFold-Inspired Framework
- Enhancing RAG Efficiency with Adaptive Context Compression
- Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
- IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
- Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
- LLM-as-a-qualitative-judge: automating error analysis in natural language generation
- Distillation versus Contrastive Learning: How to Train Your Rerankers
- FrugalRAG: Learning to retrieve and reason for multi-hop QA
- Semantic Certainty Assessment in Vector Retrieval Systems: A Novel Framework for Embedding Quality Evaluation
- Do We Really Need Specialization? Evaluating Generalist Text Embeddings for Zero-Shot Recommendation and Search
- Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization
- Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generation
- Aethorix v1.0: An Integrated Scientific AI Agent for Scalable Inorganic Materials Innovation and Industrial Implementation
- Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization
- RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
- InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
- ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
- Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
- How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
- FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
- MIRIAD: Augmenting LLMs with millions of medical query-response pairs
- CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
- Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
- GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
- Rational Retrieval Acts: Leveraging Pragmatic Reasoning to Improve Sparse Retrieval
- CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents
- Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
- CiteEval: Principle-Driven Citation Evaluation for Source Attribution
- REIC: RAG-Enhanced Intent Classification at Scale
- Rethinking Hybrid Retrieval: When Small Embeddings and LLM Re-ranking Beat Bigger Models
- From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models
- Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval
- Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
- CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents
- Ask, Retrieve, Summarize: A Modular Pipeline for Scientific Literature Summarization
- Rank-K: Test-Time Reasoning for Listwise Reranking
- An Empirical Study of Position Bias in Modern Information Retrieval
- Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
- KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval
- Incorporating Verification Standards for Security Requirements Generation from Functional Specifications
- CeQe: Grounding Lexical Retrieval in Semantic Evidence
- CRISP: Clustering Multi-Vector Representations for Denoising and Pruning
- H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
- Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction
- Improving First-stage Retrieval of Point-of-interest Search by Pre-training Models
- Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking
- Evergreen: Efficient Claim Verification for Semantic Aggregates
- PROTOCOL: Late Interaction Retrieval for Protein Homolog Search
- Revisiting Text Ranking in Deep Research
- MINT: Multi-Vector Search Index Tuning
- Formal Verification of TurboQuant: Machine-Checked Proofs and Gap Closures
- Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
- A model and package for German ColBERT
- Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers
- Credible Plan-Driven RAG Method for Multi-Hop Question Answering
- Document Optimization for Black-Box Retrieval via Reinforcement Learning
- TIFIN India at SemEval-2025: Harnessing Translation to Overcome Multilingual IR Challenges in Fact-Checked Claim Retrieval
- ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring
- FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
- EXCISE: Query-Side Exclusion for Late-Interaction Retrieval
- DynaPix: Can Vision-Language Models Identify the Exact Future?
- GraphRAFT: Retrieval Augmented Fine-Tuning for Knowledge Graphs on Graph Databases
Discussions
Related