SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
2023/10/05 by Alexander Robey, Robey, Alexander, Eric Wong +5 · 2 voices · 136 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Adversarial system #Adversary #Artificial intelligence #Attack surface #Code (set theory) #Computer science #Computer security #Conservatism #Law #Mathematics #Natural Language Processing Techniques #Point (geometry) #Political science #Programming language #Topic Modeling #Vulnerability (computing)
paper · pdf · doi:10.48550/arxiv.2310.03684
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/10/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Despite efforts to align large language models (LLMs) with human intentions, widely-used LLMs such as GPT, Llama, and Claude are susceptible to jailbreaking attacks, wherein an adversary fools a targeted LLM into generating objectionable content. To address this vulnerability, we propose SmoothLLM, the first algorithm designed to mitigate jailbreaking attacks. Based on our finding that adversarially-generated prompts are brittle to character-level changes, our defense randomly perturbs multiple copies of a given input prompt, and then aggregates the corresponding predictions to detect adversarial inputs. Across a range of popular LLMs, SmoothLLM sets the state-of-the-art for robustness against the GCG, PAIR, RandomSearch, and AmpleGCG jailbreaks. SmoothLLM is also resistant against adaptive GCG attacks, exhibits a small, though non-negligible trade-off between robustness and nominal performance, and is compatible with any LLM. Our code is publicly available at \urlhttps://github.com/arobey1/smooth-llm.
Cited by
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
- Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs
- LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- Distributional AGI Safety
- InfoFlood: Jailbreaking Large Language Models with Information Overload
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- From static to adaptive: immune memory-based jailbreak detection for large language models
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- AlignTree: Efficient Defense Against LLM Jailbreak Attacks
- AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Reimagining Safety Alignment with An Image
- Reasoning Up the Instruction Ladder for Controllable Language Models
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
- Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
- Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
- The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
- Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- Proactive defense against LLM Jailbreak
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- NonTextual Target Attack
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Takedown: How It's Done in Modern Coding Agent Exploits
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Security and Privacy Challenges of Large Language Models: A Survey
- Speculative Safety-Aware Decoding
- LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs
- Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- Semantic Representation Attack against Aligned Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- Evaluating Language Model Reasoning about Confidential Information
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Mitigating Jailbreaks with Intent-Aware LLMs
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Highlight & Summarize: RAG without the jailbreaks
- Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
- Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
- Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- A Dynamic Stackelberg Game Framework for Agentic AI Defense Against LLM Jailbreaking
- On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
- CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
- Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
- Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
- Reasoning as an Adaptive Defense for Safety
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- VERA: Variational Inference Framework for Jailbreaking Large Language Models
- Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
- Command-V: Pasting LLM Behaviors via Activation Profiles
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
- Improving Large Language Model Safety with Contrastive Representation Learning
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
- Adversarial Attacks on Robotic Vision Language Action Models
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
- Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks
- PD3F: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
- Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
- A Survey of Attacks on Large Language Models
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- Multilingual Collaborative Defense for Large Language Models
- Adversarial Suffix Filtering: a Defense Pipeline for LLMs
- One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
- Concept-Level Explainability for Auditing & Steering LLM Responses
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- Attack and defense techniques in large language models: A survey and new perspectives
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- Behavioral Integrity Verification for AI Agent Skills
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
- Frontier AI's Impact on the Cybersecurity Landscape
- DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
- DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- Geneshift: Impact of different scenario shift on Jailbreaking LLM
Discussions
Related