SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
2023/10/05 by Alexander Robey, Robey, Alexander, Eric Wong +5 · 2 voices · 80 citations
Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2310.03684
Abstract
Despite efforts to align large language models (LLMs) with human intentions, widely-used LLMs such as GPT, Llama, and Claude are susceptible to jailbreaking attacks, wherein an adversary fools a targeted LLM into generating objectionable content. To address this vulnerability, we propose SmoothLLM, the first algorithm designed to mitigate jailbreaking attacks. Based on our finding that adversarially-generated prompts are brittle to character-level changes, our defense randomly perturbs multiple copies of a given input prompt, and then aggregates the corresponding predictions to detect adversarial inputs. Across a range of popular LLMs, SmoothLLM sets the state-of-the-art for robustness against the GCG, PAIR, RandomSearch, and AmpleGCG jailbreaks. SmoothLLM is also resistant against adaptive GCG attacks, exhibits a small, though non-negligible trade-off between robustness and nominal performance, and is compatible with any LLM. Our code is publicly available at \urlhttps://github.com/arobey1/smooth-llm.
Cited by
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
- Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs
- LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- Distributional AGI Safety
- InfoFlood: Jailbreaking Large Language Models with Information Overload
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- From static to adaptive: immune memory-based jailbreak detection for large language models
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- AlignTree: Efficient Defense Against LLM Jailbreak Attacks
- AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Reimagining Safety Alignment with An Image
- Reasoning Up the Instruction Ladder for Controllable Language Models
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
- Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
- Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
- The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
- Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- Proactive defense against LLM Jailbreak
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- NonTextual Target Attack
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Takedown: How It's Done in Modern Coding Agent Exploits
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Security and Privacy Challenges of Large Language Models: A Survey
- Speculative Safety-Aware Decoding
- LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs
- Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- Semantic Representation Attack against Aligned Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- Evaluating Language Model Reasoning about Confidential Information
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Mitigating Jailbreaks with Intent-Aware LLMs
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Highlight & Summarize: RAG without the jailbreaks
Discussions
Related