Unsupervised Dense Information Retrieval with Contrastive Learning
2021/12/16 by Gautier Izacard, Izacard, Gautier, Mathilde Caron +11 · 155 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2112.09118
Abstract
Recently, information retrieval has seen the emergence of dense retrievers, using neural networks, as an alternative to classical sparse methods based on term-frequency. These models have obtained state-of-the-art results on datasets and tasks where large training sets are available. However, they do not transfer well to new applications with no training data, and are outperformed by unsupervised term-frequency methods such as BM25. In this work, we explore the limits of contrastive learning as a way to train unsupervised dense retrievers and show that it leads to strong performance in various retrieval settings. On the BEIR benchmark our unsupervised model outperforms BM25 on 11 out of 15 datasets for the Recall@100. When used as pre-training before fine-tuning, either on a few thousands in-domain examples or on the large MS~MARCO dataset, our contrastive model leads to improvements on the BEIR benchmark. Finally, we evaluate our approach for multi-lingual retrieval, where training data is even scarcer than for English, and show that our approach leads to strong unsupervised performance. Our model also exhibits strong cross-lingual transfer when fine-tuned on supervised English data only and evaluated on low resources language such as Swahili. We show that our unsupervised models can perform cross-lingual retrieval between different scripts, such as retrieving English documents from Arabic queries, which would not be possible with term matching methods.
Cited by
- TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval
- MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
- Making Large Language Models Efficient Dense Retrievers
- Mapis: A Knowledge-Graph Grounded Multi-Agent Framework for Evidence-Based PCOS Diagnosis
- Log Anomaly Detection with Large Language Models via Knowledge-Enriched Fusion
- Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- Modeling Contextual Passage Utility for Multihop Question Answering
- ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
- Factuality and Transparency Are All RAG Needs! Self-Explaining Contrastive Evidence Re-ranking
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval
- Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- Concept than Document: Context Compression via AMR-based Conceptual Entropy
- HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
- Exploring Multi-Table Retrieval Through Iterative Search
- Systematic Reconstruction of Disease Networks from Longitudinal Blood Data for Causal Discovery and Intervention Analysis
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
- Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
- ComLQ: Benchmarking Complex Logical Queries in Information Retrieval
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Search
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
- PROPEX-RAG: Enhanced GraphRAG using Prompt-Driven Prompt Execution
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Efficient Test-Time Retrieval Augmented Generation
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
- Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
- Generalized Pseudo-Relevance Feedback
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Blending Learning to Rank and Dense Representations for Efficient and Effective Cascades
- E2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
- SteerX: Disentangled Steering for LLM Personalization
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
- HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- See the Text: From Tokenization to Visual Reading
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- LLMs as Sparse Retrievers:A Framework for First-Stage Product Search
- AcademicEval: Live Long-Context LLM Benchmark
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation
- Rethinking On-policy Optimization for Query Augmentation
- OG-Rank: Learning to Rank Fast and Slow with Uncertainty and Reward-Trend Guided Adaptive Exploration
- Mixture of Experts Approaches in Dense Retrieval Tasks
- MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
- Embedding-Based Context-Aware Reranker
- Query-Specific GNN: A Comprehensive Graph Representation Learning Method for Retrieval Augmented Generation
- ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative Retrieval
- Variational Open-Domain Question Answering
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
- Study on LLMs for Promptagator-Style Dense Retriever Training
- Revisiting Long-context Modeling from Context Denoising Perspective
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- Scalable In-context Ranking with Generative Models
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- Less LLM, More Documents: Searching for Improved RAG
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- On Listwise Reranking for Corpus Feedback
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- ContextNest: Verifiable Context Governance for Autonomous AI Agent
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- Financial Risk Relation Identification through Dual-view Adaptation
- Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- Towards Effective and Efficient Sparse Neural Information Retrieval
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- CORE-RAG: Lossless Compression for Retrieval-Augmented LLMs via Reinforcement Learning
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Boosting Data Utilization for Multilingual Dense Retrieval
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
- HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
- Upcycling Candidate Tokens of Large Language Models for Query Expansion
- ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
- Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
- QZhou-Embedding Technical Report
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Ontology-Guided Query Expansion for Biomedical Document Retrieval using Large Language Models
- Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering
- Learning from Natural Language Feedback for Personalized Question Answering
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval
- SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
- Synthesizing scientific literature with retrieval-augmented language models
- Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation
- FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
- Knowledge Editing for Multi-Hop Question Answering Using Semantic Analysis
- Latent Inter-User Difference Modeling for LLM Personalization
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Customize Multi-modal RAI Guardrails with Precedent-based predictions
- A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
- Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
Related