2020/10/15 by Zijie J. Wang, Zijie Wang, Lixi Zhou +10
Computer Science · Decision Sciences · #Anomaly Detection Techniques and Applications #Data Quality and Management #Data Stream Mining Techniques #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.DB #cs.LG
paper · pdf · doi:10.48550/arxiv.2010.07586
In submission
arxiv created 2020/10/15 · openalex publication_date 2020/10/15 · arxiv updated 2020/10/16 · openalex created_date 2020/10/22 · openalex updated_date 2026/07/28
Data is the king in the age of AI. However data integration is often a laborious task that is hard to automate. Schema change is one significant obstacle to the automation of the end-to-end data integration process. Although there exist mechanisms such as query discovery and schema modification language to handle the problem, these approaches can only work with the assumption that the schema is maintained by a database. However, we observe diversified schema changes in heterogeneous data and open data, most of which has no schema defined. In this work, we propose to use deep learning to automatically deal with schema changes through a super cell representation and automatic injection of perturbations to the training data to make the model robust to schema changes. Our experimental results demonstrate that our proposed approach is effective for two real-world data integration scenarios: coronavirus data integration, and machine log integration.