A Privacy-Preserving Approach to Extraction of Personal Information through Automatic Annotation and Federated Learning

We curated WikiPII, an automatically labeled dataset composed of Wikipedia\nbiography pages, annotated for personal information extraction. Although\nautomatic annotation can lead to a high degree of label noise, it is an\ninexpensive process and can generate large volumes of annotated documents. We\ntrained a BERT-based NER model with WikiPII and showed that with an adequately\nlarge training dataset, the model can significantly decrease the cost of manual\ninformation extraction, despite the high level of label noise. In a similar\napproach, organizations can leverage text mining techniques to create\ncustomized annotated datasets from their historical data without sharing the\nraw data for human annotation. Also, we explore collaborative training of NER\nmodels through federated learning when the annotation is noisy. Our results\nsuggest that depending on the level of trust to the ML operator and the volume\nof the available data, distributed training can be an effective way of training\na personal information identifier in a privacy-preserved manner. Research\nmaterial is available at https://github.com/ratmcu/wikipiifed.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC