JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
2024/03/28 by Patrick Chao, Edoardo Debenedetti, Chao, Patrick +21 · 180 citations
Computer Science · #Authorship Attribution and Profiling #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2404.01318
openalex publication_date 2024/03/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Jailbreak attacks cause large language models (LLMs) to generate harmful, unethical, or otherwise objectionable content. Evaluating these attacks presents a number of challenges, which the current collection of benchmarks and evaluation techniques do not adequately address. First, there is no clear standard of practice regarding jailbreaking evaluation. Second, existing works compute costs and success rates in incomparable ways. And third, numerous works are not reproducible, as they withhold adversarial prompts, involve closed-source code, or rely on evolving proprietary APIs. To address these challenges, we introduce JailbreakBench, an open-sourced benchmark with the following components: (1) an evolving repository of state-of-the-art adversarial prompts, which we refer to as jailbreak artifacts; (2) a jailbreaking dataset comprising 100 behaviors -- both original and sourced from prior work (Zou et al., 2023; Mazeika et al., 2023, 2024) -- which align with OpenAI's usage policies; (3) a standardized evaluation framework at https://github.com/JailbreakBench/jailbreakbench that includes a clearly defined threat model, system prompts, chat templates, and scoring functions; and (4) a leaderboard at https://jailbreakbench.github.io/ that tracks the performance of attacks and defenses for various LLMs. We have carefully considered the potential ethical implications of releasing this benchmark, and believe that it will be a net positive for the community.
Cited by
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- Eliciting Behaviors in Multi-Turn Conversations
- ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- LLM Scheming Inversely Scales with Pretraining Language Coverage
- DREAM: Dynamic Red-teaming across Environments for AI Models
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Prompt Governance? On Governing Technologies Governed by Natural Language
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
- Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- Automating Deception: Scalable Multi-Turn LLM Jailbreaks
- ASTRA: Agentic Steerability and Risk Assessment Framework
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Efficient LLM Safety Evaluation through Multi-Agent Debate
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Diffusion LLMs are Natural Adversaries for any LLM
- Chain-of-Thought Hijacking
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- PromptLocate: Localizing Prompt Injection Attacks
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- When to Reason: Semantic Router for vLLM
- Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
- LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Proactive defense against LLM Jailbreak
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- NonTextual Target Attack
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- A Granular Study of Safety Pretraining under Model Abliteration
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
- RobustFlow: Towards Robust Agentic Workflow Generation
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- Evaluating Language Model Reasoning about Confidential Information
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Mitigating Jailbreaks with Intent-Aware LLMs
- Jinx: Unlimited LLMs for Probing Alignment Failures
- Multi-Turn Jailbreaks Are Simpler Than They Seem
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- Measuring Harmfulness of Computer-Using Agents
- The Blessing and Curse of Dimensionality in Safety Alignment
- PurpCode: Reasoning for Safer Code Generation
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- LLMs Encode Harmfulness and Refusal Separately
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
- Medical Red Teaming Protocol of Language Models: On the Importance of User Perspectives in Healthcare Settings
- Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
- The bitter lesson of misuse detection
- Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
- Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- TAI3: Testing Agent Integrity in Interpreting User Intent
- Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
- InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- LLM Jailbreak Oracle
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
- Universal Jailbreak Suffixes Are Strong Attention Hijackers
- SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- Exploring the Secondary Risks of Large Language Models
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Should LLM Safety Be More Than Refusing Harmful Instructions?
- Adversarial Attacks on Robotic Vision Language Action Models
- HASHIRU: Hierarchical Agent System for Hybrid Intelligent Resource Utilization
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- Learning Safety Constraints for Large Language Models
- Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
- MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
- GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- Concealment of Intent: A Game-Theoretic Analysis
- SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
- The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
- Mitigating Deceptive Alignment via Self-Monitoring
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
- Robustifying Vision-Language Models via Dynamic Token Reweighting
- Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
- Advancing LLM Safe Alignment with Safety Representation Ranking
- Causes and Consequences of Representational Similarity in Machine Learning Models
- SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- A Survey of Attacks on Large Language Models
- Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
- Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
- RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
- When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
- Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
- Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement
- DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in LLMs
- Intent Laundering: AI Safety Datasets Are Not What They Seem
- XBreaking: Understanding how LLMs security alignment can be broken
- Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
- Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
- Targeted Neuron Modulation via Contrastive Pair Search
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
- Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing
- Steered LLM Activations are Non-Surjective
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- Security Steerability is All You Need
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- How Should AI Safety Benchmarks Benchmark Safety?
- Safety Pretraining: Toward the Next Generation of Safe AI
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- Detecting Safety Training Modification in Language Models via Activation Analysis
- RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
- The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Related