vix.ing · top · new · best · stats · spec

Identifying Untrustworthy Predictions in Neural Networks by Geometric\n Gradient Analysis

2021/02/24 by Leo Schwinn, An Nguyen, Schwinn, Leo +13 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications

paper · pdf · doi:10.48550/arxiv.2102.12196

Abstract

The susceptibility of deep neural networks to untrustworthy predictions,\nincluding out-of-distribution (OOD) data and adversarial examples, still\nprevent their widespread use in safety-critical applications. Most existing\nmethods either require a re-training of a given model to achieve robust\nidentification of adversarial attacks or are limited to out-of-distribution\nsample detection only. In this work, we propose a geometric gradient analysis\n(GGA) to improve the identification of untrustworthy predictions without\nretraining of a given model. GGA analyzes the geometry of the loss landscape of\nneural networks based on the saliency maps of their respective input. To\nmotivate the proposed approach, we provide theoretical connections between\ngradients' geometrical properties and local minima of the loss function.\nFurthermore, we demonstrate that the proposed method outperforms prior\napproaches in detecting OOD data and adversarial attacks, including\nstate-of-the-art and adaptive attacks.\n

Citations

Cited by

Related