CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
2025/04/15 by Zhong, Hanmeng, Chen, Linqing, Wu, Wentao +1
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2504.12342
Abstract
Recent development in Retrieval-Augmented Large Language Models (LLMs) have shown great promise in biomedical applications. How ever, a critical gap persists in reliably evaluating their curation ability the process by which models select and integrate relevant references while filtering out noise. To address this, we introduce the benchmark for Curation of Retrieval-Augmented LLMs in Biomedicine (CRAB), the first multilingual benchmark tailored for evaluating the biomedical curation of retrieval-augmented LLMs, available in English, French, German and Chinese. By incorporating a novel citation-based evaluation metric, CRAB quantifies the curation performance of retrieval-augmented LLMs in biomedicine. Experimental results reveal significant discrepancies in the curation performance of mainstream LLMs, underscoring the urgent need to improve it in the domain of biomedicine. Our dataset is available at https://huggingface.co/datasets/zhm0/CRAB.
Citations
- Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
- Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
- RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
- The Llama 3 Herd of Models
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Benchmarking Retrieval-Augmented Generation for Medicine
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
- Design and Evaluation of a Retrieval-Augmented Generation Architecture for OWASP Security Data
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- Lost in the Middle: How Language Models Use Long Contexts
- Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset
- Enabling Large Language Models to Generate Text with Citations
- Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks
- Capabilities of GPT-4 on Medical Challenge Problems
- Almanac: Retrieval-Augmented Language Models for Clinical Medicine
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- Rethinking with Retrieval: Faithful Large Language Model Inference
- Large Language Models Encode Clinical Knowledge
- Large language models encode clinical knowledge
- Large Language Models with Controllable Working Memory
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- Teaching language models to support answers with verified quotes
- Improving language models by retrieving from trillions of tokens
- The Curious Case of Hallucinations in Neural Machine Translation
- Factual Error Correction for Abstractive Summarization Models
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- Measuring Massive Multitask Language Understanding
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- REALM: Retrieval-Augmented Language Model Pre-Training
- PubMedQA: A Dataset for Biomedical Research Question Answering
Related