vix.ing · top · new · best · stats · spec

A Privacy-Preserving Approach to Extraction of Personal Information\n through Automatic Annotation and Federated Learning

2021/05/19 by Rajitha Hathurusinghe, Hathurusinghe, Rajitha, Isar Nejadgholi +3 · 2 citations
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #Data Quality and Management #Digital and Cyber Forensics #FOS: Computer and information sciences #Privacy-Preserving Technologies in Data

paper · pdf · doi:10.48550/arxiv.2105.09198

openalex publication_date 2021/05/19 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

We curated WikiPII, an automatically labeled dataset composed of Wikipedia\nbiography pages, annotated for personal information extraction. Although\nautomatic annotation can lead to a high degree of label noise, it is an\ninexpensive process and can generate large volumes of annotated documents. We\ntrained a BERT-based NER model with WikiPII and showed that with an adequately\nlarge training dataset, the model can significantly decrease the cost of manual\ninformation extraction, despite the high level of label noise. In a similar\napproach, organizations can leverage text mining techniques to create\ncustomized annotated datasets from their historical data without sharing the\nraw data for human annotation. Also, we explore collaborative training of NER\nmodels through federated learning when the annotation is noisy. Our results\nsuggest that depending on the level of trust to the ML operator and the volume\nof the available data, distributed training can be an effective way of training\na personal information identifier in a privacy-preserved manner. Research\nmaterial is available at https://github.com/ratmcu/wikipiifed.\n

Cited by

Related