2019/11/15 by Sean Augenstein, H. Brendan McMahan, Augenstein, Sean +14 · 43 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Artificial intelligence #Computer science #Data mining #Debugging #Differential privacy #Enhanced Data Rates for GSM Evolution #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Federated learning #Generative grammar #Generative model #Intuition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Outlier #Privacy-Preserving Technologies in Data #Raw data #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1911.06679
published in arXiv (Cornell University) (Cornell University) · 26 pages, 8 figures. Camera-ready ICLR 2020 version
openalex publication_date 2019/11/15 · arxiv created 2020/02/04 · arxiv updated 2020/02/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
To improve real-world applications of machine learning, experienced modelers develop intuition about their datasets, their models, and how the two interact. Manual inspection of raw data - of representative samples, of outliers, of misclassifications - is an essential tool in a) identifying and fixing problems in the data, b) generating new modeling hypotheses, and c) assigning or refining human-provided labels. However, manual data inspection is problematic for privacy sensitive datasets, such as those representing the behavior of real-world individuals. Furthermore, manual data inspection is impossible in the increasingly important setting of federated learning, where raw examples are stored at the edge and the modeler may only access aggregated outputs such as metrics or model parameters. This paper demonstrates that generative models - trained using federated methods and with formal differential privacy guarantees - can be used effectively to debug many commonly occurring data issues even when the data cannot be directly inspected. We explore these methods in applications to text with differentially private federated RNNs and to images using a novel algorithm for differentially private federated GANs.