2025/11/20 by Marijan Soric, Soric, Marijan, Cécile Gracianne +5
Computer Science · Decision Sciences · #Benchmark (surveying) #Benchmarking #Data Quality and Management #Databases (cs.DB) #FOS: Computer and information sciences #Generalizability theory #Handwritten Text Recognition Techniques #Robustness (evolution) #Software #Table (database) #Web Data Mining and Analysis #cs.DB
paper · pdf · doi:10.48550/arxiv.2511.16134
published in arXiv (Cornell University) (Cornell University) · Extended version, with appendices, of a paper published at KDD '26
openalex publication_date 2025/11/20 · openalex created_date 2025/11/23 · arxiv created 2026/07/31 · arxiv updated 2026/08/03 · openalex updated_date 2026/08/05
Table Extraction (TE) consists in extracting tables from PDF documents, in a structured format enabling automatic processing. While numerous TE tools exist, the variety of methods and techniques makes it difficult for users to choose the most appropriate one. We propose a novel benchmark for assessing end-to-end TE methods (from PDF to the final table) over 86k pages. We contribute an analysis of TE evaluation metrics, and a novel, rigorous evaluation process, which allows scoring each TE sub-task as well as end-to-end TE, and captures model uncertainty. Along with prior datasets, our benchmark comprises two new heterogeneous datasets of 39k samples. We run our benchmark on diverse models, including off-the-shelf libraries, tools, computer vision-based models and modern approaches using general and specialized vision language models. The results demonstrate that TE remains challenging: current methods suffer from a lack of generalizability when facing heterogeneous data, and from limitations in robustness and interpretability.