AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
2023/10/03 by Xiaogeng Liu, Liu, Xiaogeng, Nan Xu +5 · 259 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2310.04451
openalex publication_date 2023/10/03 · openalex created_date 2023/10/12 · openalex updated_date 2026/07/28
Abstract
The aligned Large Language Models (LLMs) are powerful language understanding and decision-making tools that are created through extensive alignment with human feedback. However, these large models remain susceptible to jailbreak attacks, where adversaries manipulate prompts to elicit malicious outputs that should not be given by aligned LLMs. Investigating jailbreak prompts can lead us to delve into the limitations of LLMs and further guide us to secure them. Unfortunately, existing jailbreak techniques suffer from either (1) scalability issues, where attacks heavily rely on manual crafting of prompts, or (2) stealthiness problems, as attacks depend on token-based algorithms to generate prompts that are often semantically meaningless, making them susceptible to detection through basic perplexity testing. In light of these challenges, we intend to answer this question: Can we develop an approach that can automatically generate stealthy jailbreak prompts? In this paper, we introduce AutoDAN, a novel jailbreak attack against aligned LLMs. AutoDAN can automatically generate stealthy jailbreak prompts by the carefully designed hierarchical genetic algorithm. Extensive evaluations demonstrate that AutoDAN not only automates the process while preserving semantic meaningfulness, but also demonstrates superior attack strength in cross-model transferability, and cross-sample universality compared with the baseline. Moreover, we also compare AutoDAN with perplexity-based defense methods and show that AutoDAN can bypass them effectively.
Cited by
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- Eliciting Behaviors in Multi-Turn Conversations
- Agent Data Injection Attacks are Realistic Threats to AI Agents
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
- In-Context Representation Hijacking
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
- Securing Large Language Models (LLMs) from Prompt Injection Attacks
- DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Evolving Prompts for Toxicity Search in Large Language Models
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- "Power of Words": Stealthy and Adaptive Private Information Elicitation via LLM Communication Strategies
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Efficient LLM Safety Evaluation through Multi-Agent Debate
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Jailbreaking in the Haystack
- Reimagining Safety Alignment with An Image
- Diffusion LLMs are Natural Adversaries for any LLM
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- Evolvable AI: Threats of a new major transition in evolution
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
- Soft Instruction De-escalation Defense
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- Black-box Optimization of LLM Outputs by Asking for Directions
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
- TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Deep Research Brings Deeper Harm
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- Inverse Language Modeling towards Robust and Grounded LLMs
- PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
- MCCE: A Framework for Multi-LLM Collaborative Co-Evolution
- Proactive defense against LLM Jailbreak
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- NonTextual Target Attack
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Model Correlation Detection via Random Selection Probing
- RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- GuardNet: Graph-Attention Filtering for Jailbreak Defense in Large Language Models
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- Security and Privacy Challenges of Large Language Models: A Survey
- Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
- Semantic Representation Attack against Aligned Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- GoldenTransformer: A Modular Fault Injection Framework for Transformer Robustness Research
- Steering MoE LLMs via Expert (De)Activation
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- Reliable Weak-to-Strong Monitoring of LLM Agents
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- Involuntary Jailbreak: On Self-Prompting Attacks
- Mitigating Jailbreaks with Intent-Aware LLMs
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment
- Special-Character Adversarial Attacks on Open-Source Language Model
- A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
- Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- Enhancing Jailbreak Attacks on LLMs via Persona Prompts
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- PurpCode: Reasoning for Safer Code Generation
- The Geometry of Harmfulness in LLMs through Subconcept Probing
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
- On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
- Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
- PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
- Reasoning as an Adaptive Defense for Safety
- Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
- VERA: Variational Inference Framework for Jailbreaking Large Language Models
- TAI3: Testing Agent Integrity in Interpreting User Intent
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
- Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- Probing the Robustness of Large Language Models Safety to Latent Perturbations
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- LLM Jailbreak Oracle
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- ContextBench: Modifying Contexts for Targeted Latent Activation
- Universal Jailbreak Suffixes Are Strong Attention Hijackers
- SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
- Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
- TracLLM: A Generic Framework for Attributing Long Context LLMs
- Should LLM Safety Be More Than Refusing Harmful Instructions?
- Adversarial Attacks on Robotic Vision Language Action Models
- Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- Learning Safety Constraints for Large Language Models
- LLM Agents Should Employ Security Principles
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- Does Machine Unlearning Truly Remove Knowledge?
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
- GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- System Prompt Extraction Attacks and Defenses in Large Language Models
- Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
- Lifelong Safety Alignment for Language Models
- Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
- Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
- Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
- Robustifying Vision-Language Models via Dynamic Token Reweighting
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
- Adversarially Pretrained Transformers may be Universally Robust In-Context Learners
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Safety Alignment Can Be Not Superficial With Explicit Safety Signals
- Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
- JULI: Jailbreak Large Language Models by Self-Introspection
- Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
- Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
- Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
- Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers
- Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment
- Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
- OET: Optimization-based prompt injection Evaluation Toolkit
- Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement
- XBreaking: Understanding how LLMs security alignment can be broken
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
- Do Thinking Tokens Help with Safety?
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- The Automation Advantage in AI Red Teaming
- Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
- Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
- Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
- Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
- A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
- RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
- DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
- Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
- GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
- Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
- The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
- RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- Geneshift: Impact of different scenario shift on Jailbreaking LLM
Related