Jailbreaking Black Box Large Language Models in Twenty Queries
2023/10/12 by Patrick Chao, Chao, Patrick, Alexander Robey +9 · 216 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2310.08419
openalex publication_date 2023/10/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
There is growing interest in ensuring that large language models (LLMs) align with human values. However, the alignment of such models is vulnerable to adversarial jailbreaks, which coax LLMs into overriding their safety guardrails. The identification of these vulnerabilities is therefore instrumental in understanding inherent weaknesses and preventing future misuse. To this end, we propose Prompt Automatic Iterative Refinement (PAIR), an algorithm that generates semantic jailbreaks with only black-box access to an LLM. PAIR -- which is inspired by social engineering attacks -- uses an attacker LLM to automatically generate jailbreaks for a separate targeted LLM without human intervention. In this way, the attacker LLM iteratively queries the target LLM to update and refine a candidate jailbreak. Empirically, PAIR often requires fewer than twenty queries to produce a jailbreak, which is orders of magnitude more efficient than existing algorithms. PAIR also achieves competitive jailbreaking success rates and transferability on open and closed-source LLMs, including GPT-3.5/4, Vicuna, and Gemini.
Cited by
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- Boundary Point Jailbreaking of Black-Box LLMs
- Do LLMs Know Their Vulnerable Scenarios?
- TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- AgentWatcher: A Rule-based Prompt Injection Monitor
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- Safety Alignment of LMs via Non-cooperative Games
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
- Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
- Are Your Agents Upward Deceivers?
- In-Context Representation Hijacking
- From static to adaptive: immune memory-based jailbreak detection for large language models
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- An Invariant Latent Space Perspective on Language Model Inversion
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- Robustness of LLM-enabled vehicle trajectory prediction under data security threats
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- Efficient LLM Safety Evaluation through Multi-Agent Debate
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
- Jailbreaking in the Haystack
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
- Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
- Reimagining Safety Alignment with An Image
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- Diffusion LLMs are Natural Adversaries for any LLM
- Reasoning Up the Instruction Ladder for Controllable Language Models
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- A Survey on Unlearning in Large Language Models
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
- Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
- Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- Black-box Optimization of LLM Outputs by Asking for Directions
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
- On the Convergence of Moral Self-Correction in Large Language Models
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
- Learning to Interpret Weight Differences in Language Models
- Imperceptible Jailbreaking against Large Language Models
- Proactive defense against LLM Jailbreak
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
- NonTextual Target Attack
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- CHAI: Command Hijacking against embodied AI
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
- STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
- SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
- Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- AI Agents May Always Fall for Prompt Injections
- Security and Privacy Challenges of Large Language Models: A Survey
- Adversarial machine learning :
- Speculative Safety-Aware Decoding
- Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- Semantic Representation Attack against Aligned Large Language Models
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- MillStone: How Open-Minded Are LLMs?
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- Evaluating Language Model Reasoning about Confidential Information
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
- Mitigating Jailbreaks with Intent-Aware LLMs
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Many-Turn Jailbreaking
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks
- JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
- Automatic LLM Red Teaming
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
- Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- Activation-Guided Local Editing for Jailbreaking Attacks
- Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
- Strategic Deflection: Defending LLMs from Logit Manipulation
Related