2024/02/27 by Vyas Raina, Samson Tan, Raina, Vyas +9 · 1 citation
Arts and Humanities · Biochemistry, Genetics and Molecular Biology · Social Sciences · #Bacillus and Francisella bacterial research #Computation and Language (cs.CL) #FOS: Computer and information sciences #History of Science and Medicine #Intelligence, Security, War Strategy
paper · pdf · doi:10.48550/arxiv.2402.17509
openalex publication_date 2024/02/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Deep learning-based Natural Language Processing (NLP) models are vulnerable to adversarial attacks, where small perturbations can cause a model to misclassify. Adversarial Training (AT) is often used to increase model robustness. However, we have discovered an intriguing phenomenon: deliberately or accidentally miscalibrating models masks gradients in a way that interferes with adversarial attack search methods, giving rise to an apparent increase in robustness. We show that this observed gain in robustness is an illusion of robustness (IOR), and demonstrate how an adversary can perform various forms of test-time temperature calibration to nullify the aforementioned interference and allow the adversarial attack to find adversarial examples. Hence, we urge the NLP community to incorporate test-time temperature scaling into their robustness evaluations to ensure that any observed gains are genuine. Finally, we show how the temperature can be scaled during training to improve genuine robustness.