vix.ing · top · new · best · stats

Position: Adversarial ML for LLMs Is Not Making Any Progress

2025/02/04 by Javier Rando, Rando, Javier, Jie Zhang +5 · 3 voices · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CR #cs.LG

paper · pdf · doi:10.48550/arxiv.2502.02260

openalex publication_date 2025/02/04 · arxiv published 2025/02/04 · openalex created_date 2025/10/10 · arxiv updated 2026/06/02 · openalex updated_date 2026/07/28

Abstract

In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even for simple "toy" problems (e.g., robustness to small adversarial perturbations) and is often hindered by non-rigorous evaluations. Today, adversarial ML research has shifted towards studying larger, general-purpose language models. In this position paper, we argue that the situation is now even worse: in the era of LLMs, the field of adversarial ML studies problems that are (1) less clearly defined, (2) harder to solve, and (3) even more challenging to evaluate. As a result, we caution that yet another decade of work on adversarial ML may be failing to produce meaningful progress.

Cited by

Discussions

Related