vix.ing · top · new · best · stats · spec

ORB: An Open Reading Benchmark for Comprehensive Evaluation of Machine\n Reading Comprehension

2019/12/29 by Dheeru Dua, Ananth Gottumukkala, Dua, Dheeru +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1912.12598

openalex publication_date 2019/12/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Reading comprehension is one of the crucial tasks for furthering research in\nnatural language understanding. A lot of diverse reading comprehension datasets\nhave recently been introduced to study various phenomena in natural language,\nranging from simple paraphrase matching and entity typing to entity tracking\nand understanding the implications of the context. Given the availability of\nmany such datasets, comprehensive and reliable evaluation is tedious and\ntime-consuming for researchers working on this problem. We present an\nevaluation server, ORB, that reports performance on seven diverse reading\ncomprehension datasets, encouraging and facilitating testing a single model's\ncapability in understanding a wide variety of reading phenomena. The evaluation\nserver places no restrictions on how models are trained, so it is a suitable\ntest bed for exploring training paradigms and representation learning for\ngeneral reading facility. As more suitable datasets are released, they will be\nadded to the evaluation server. We also collect and include synthetic\naugmentations for these datasets, testing how well models can handle\nout-of-domain questions.\n

Citations

Cited by

Related