vix.ing · top · new · best · stats

LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation

2025/11/09 by Liya Zhu, Zhu, Liya, Peizhuang Cong +44 · 1 voice · 1 citation
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL

paper · pdf · doi:10.48550/arxiv.2511.06346

openalex publication_date 2025/11/09 · arxiv published 2025/11/09 · openalex created_date 2025/11/12 · arxiv updated 2026/01/08 · openalex updated_date 2026/07/28

Abstract

Large Language Models (LLMs) perform well on standard reasoning and question-answering benchmarks, yet such evaluations often fail to capture their ability to handle long-tail, expertise-intensive knowledge in real-world professional scenarios. We introduce LPFQA, a long-tail knowledge benchmark derived from authentic professional forum discussions, covering 7 academic and industrial domains with 430 curated tasks grounded in practical expertise. LPFQA evaluates specialized reasoning, domain-specific terminology understanding, and contextual interpretation, and adopts a hierarchical difficulty structure to ensure semantic clarity and uniquely identifiable answers. Experiments on over multiple mainstream LLMs reveal substantial performance gaps, particularly on tasks requiring deep domain reasoning, exposing limitations overlooked by existing benchmarks. Overall, LPFQA provides an authentic and discriminative evaluation framework that complements prior benchmarks and informs future LLM development.

Citations

Cited by

Discussions

Related