vix.ing · top · new · best · stats · spec

IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review

2026/04/23 by Fengbo Ma, Zixin Rao, Xiaoting Li +5 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Information Retrieval and Search Behavior #Topic Modeling #cs.AI #cs.IR #cs.LG

paper · pdf · doi:10.48550/arxiv.2604.22861

openalex publication_date 2026/04/23 · arxiv published 2026/04/23 · arxiv updated 2026/04/23 · openalex created_date 2026/04/29 · openalex updated_date 2026/07/28

Abstract

Scientific research relies on accurate information retrieval from literature to support analytical decisions. In this work, we introduce a new task, INformation reTRieval through literAture reVIEW (IntraView), which aims to automate fine-grained information retrieval faithfully grounded in the provided content in response to research-driven queries, and propose IntrAgent, an LLM-based agent that addresses this challenging task. In particular, IntrAgent is designed to mimic human behaviors when reading literature for information retrieval -- identifying relevant sections and then iteratively extracting key details to refine the retrieved information. It follows a two-stage pipeline: a Section Ranking stage that prioritizes relevant literature sections through structural-knowledge-enabled reasoning, and an Iterative Reading stage that continuously extracts details and synthesizes them into concise, contextually grounded answers. To support rigorous evaluation, we introduce IntraBench, a new benchmark consisting of 315 test instances built from expert-authored questions paired with literature spanning five STEM domains. Across seven backbone LLMs, IntrAgent achieves on average 13.2% higher cross-domain accuracy than state-of-the-art RAG and research-agent baselines.

Citations

Discussions