2020/11/05 by Michał J. Gajda, Hai Nguyen Quang, Gajda, Michał J. +5
Computer Science · Decision Sciences · #Advanced Database Systems and Queries #FOS: Computer and information sciences #Programming Languages (cs.PL) #Scientific Computing and Data Management #Web Data Mining and Analysis #cs.PL
paper · pdf · doi:10.48550/arxiv.2011.03538
arxiv created 2020/11/05 · openalex publication_date 2020/11/05 · arxiv updated 2020/11/10 · openalex created_date 2024/04/10 · openalex updated_date 2026/07/28
We propose reformulation of discovery of data structure within a web page as relations between sets of document nodes. We start by reformulating web page analysis as finding expressions in extension of XPath. Then we propose to automatically discover these XPath expressions with InferXPath meta-language. Our goal is to automate laborious process of conversion of manually created web pages that serve as software documentations, wikis, and reference documents, and speed up their conversion into tabular data that can be directly fed into data pipeline.