2025/11/15 by Thong Bach, Dung Nguyen, Bach, Thong +5
Computer Science · #Adversarial Robustness in Machine Learning #Adversarial system #Autoregressive model #Deep learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Language model #Machine Learning (cs.LG) #Task analysis #Topic Modeling #Vocabulary
paper · pdf · doi:10.48550/arxiv.2511.12155
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/11/15 · openalex created_date 2025/11/19 · openalex updated_date 2026/08/05
Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment. We provide a mechanistic analysis revealing that position-dependent gradient weakening during autoregressive training creates signal decay, leading to incomplete safety learning where safety training fails to transform model preferences in later response regions fully. We introduce base-favored tokens -- vocabulary elements where base models assign higher probability than aligned models -- as computational indicators of incomplete safety learning and develop a targeted completion method that addresses undertrained regions through adaptive penalties and hybrid teacher distillation. Experimental evaluation across Llama and Qwen model families demonstrates dramatic improvements in adversarial robustness, with 48--98% reductions in attack success rates while preserving general capabilities. These results establish both a mechanistic understanding and practical solutions for fundamental limitations in safety alignment methodologies.