2021/04/05 by Abdelghny Orogat, Orogat, Abdelghny, Isabelle Liu +3 · 1 citation
Computer Science · #Advanced Graph Neural Networks #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2105.00811
openalex publication_date 2021/04/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recently, there has been an increase in the number of knowledge graphs that\ncan be only queried by experts. However, describing questions using structured\nqueries is not straightforward for non-expert users who need to have sufficient\nknowledge about both the vocabulary and the structure of the queried knowledge\ngraph, as well as the syntax of the structured query language used to describe\nthe user's information needs. The most popular approach introduced to overcome\nthe aforementioned challenges is to use natural language to query these\nknowledge graphs. Although several question answering benchmarks can be used to\nevaluate question-answering systems over a number of popular knowledge graphs,\nchoosing a benchmark to accurately assess the quality of a question answering\nsystem is a challenging task.\n In this paper, we introduce CBench, an extensible, and more informative\nbenchmarking suite for analyzing benchmarks and evaluating question answering\nsystems. CBench can be used to analyze existing benchmarks with respect to\nseveral fine-grained linguistic, syntactic, and structural properties of the\nquestions and queries in the benchmark. We show that existing benchmarks vary\nsignificantly with respect to these properties deeming choosing a small subset\nof them unreliable in evaluating QA systems. Until further research improves\nthe quality and comprehensiveness of benchmarks, CBench can be used to\nfacilitate this evaluation using a set of popular benchmarks that can be\naugmented with other user-provided benchmarks. CBench not only evaluates a\nquestion answering system based on popular single-number metrics but also gives\na detailed analysis of the linguistic, syntactic, and structural properties of\nanswered and unanswered questions to better help the developers of question\nanswering systems to better understand where their system excels and where it\nstruggles.\n