2020/01/28 by Farhad Farokhi, Farokhi, Farhad · 1 citation
Mathematics · #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.2001.10655
We use distributionally-robust optimization for machine learning to mitigate\nthe effect of data poisoning attacks. We provide performance guarantees for the\ntrained model on the original data (not including the poison records) by\ntraining the model for the worst-case distribution on a neighbourhood around\nthe empirical distribution (extracted from the training dataset corrupted by a\npoisoning attack) defined using the Wasserstein distance. We relax the\ndistributionally-robust machine learning problem by finding an upper bound for\nthe worst-case fitness based on the empirical sampled-averaged fitness and the\nLipschitz-constant of the fitness function (on the data for given model\nparameters) as regularizer. For regression models, we prove that this\nregularizer is equal to the dual norm of the model parameters. We use the Wine\nQuality dataset, the Boston Housing Market dataset, and the Adult dataset for\ndemonstrating the results of this paper.\n