vix.ing · top · new · best · stats · spec

Gradient-based Adversarial Attacks against Text Transformers

2021/04/15 by Chuan Guo, Guo, Chuan, Alexandre Sablayrolles +5 · 18 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2104.13733

openalex publication_date 2021/04/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.

Citations

Cited by

Related