2023/07/06 by Siddharth Khincha, Khincha, Siddharth, Chelsi Jain +7 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Data Quality and Management #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2307.03313
openalex publication_date 2023/07/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Information Synchronization of semi-structured data across languages is challenging. For instance, Wikipedia tables in one language should be synchronized across languages. To address this problem, we introduce a new dataset InfoSyncC and a two-step method for tabular synchronization. InfoSync contains 100K entity-centric tables (Wikipedia Infoboxes) across 14 languages, of which a subset (3.5K pairs) are manually annotated. The proposed method includes 1) Information Alignment to map rows and 2) Information Update for updating missing/outdated information for aligned tables across multilingual tables. When evaluated on InfoSync, information alignment achieves an F1 score of 87.91 (en <-> non-en). To evaluate information updation, we perform human-assisted Wikipedia edits on Infoboxes for 603 table pairs. Our approach obtains an acceptance rate of 77.28% on Wikipedia, showing the effectiveness of the proposed method.