JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
2024/03/28 by Patrick Chao, Edoardo Debenedetti, Chao, Patrick +21 · 93 citations
Computer Science · #Authorship Attribution and Profiling #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2404.01318
openalex publication_date 2024/03/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Jailbreak attacks cause large language models (LLMs) to generate harmful, unethical, or otherwise objectionable content. Evaluating these attacks presents a number of challenges, which the current collection of benchmarks and evaluation techniques do not adequately address. First, there is no clear standard of practice regarding jailbreaking evaluation. Second, existing works compute costs and success rates in incomparable ways. And third, numerous works are not reproducible, as they withhold adversarial prompts, involve closed-source code, or rely on evolving proprietary APIs. To address these challenges, we introduce JailbreakBench, an open-sourced benchmark with the following components: (1) an evolving repository of state-of-the-art adversarial prompts, which we refer to as jailbreak artifacts; (2) a jailbreaking dataset comprising 100 behaviors -- both original and sourced from prior work (Zou et al., 2023; Mazeika et al., 2023, 2024) -- which align with OpenAI's usage policies; (3) a standardized evaluation framework at https://github.com/JailbreakBench/jailbreakbench that includes a clearly defined threat model, system prompts, chat templates, and scoring functions; and (4) a leaderboard at https://jailbreakbench.github.io/ that tracks the performance of attacks and defenses for various LLMs. We have carefully considered the potential ethical implications of releasing this benchmark, and believe that it will be a net positive for the community.
Cited by
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- Eliciting Behaviors in Multi-Turn Conversations
- ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- LLM Scheming Inversely Scales with Pretraining Language Coverage
- DREAM: Dynamic Red-teaming across Environments for AI Models
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Prompt Governance? On Governing Technologies Governed by Natural Language
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Challenges of Evaluating LLM Safety for User Welfare
- Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- Automating Deception: Scalable Multi-Turn LLM Jailbreaks
- ASTRA: Agentic Steerability and Risk Assessment Framework
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Efficient LLM Safety Evaluation through Multi-Agent Debate
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Diffusion LLMs are Natural Adversaries for any LLM
- Chain-of-Thought Hijacking
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- PromptLocate: Localizing Prompt Injection Attacks
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- When to Reason: Semantic Router for vLLM
- Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
- LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Proactive defense against LLM Jailbreak
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- NonTextual Target Attack
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- A Granular Study of Safety Pretraining under Model Abliteration
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
- RobustFlow: Towards Robust Agentic Workflow Generation
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- Evaluating Language Model Reasoning about Confidential Information
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Mitigating Jailbreaks with Intent-Aware LLMs
- Jinx: Unlimited LLMs for Probing Alignment Failures
- Multi-Turn Jailbreaks Are Simpler Than They Seem
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- Measuring Harmfulness of Computer-Using Agents
- The Blessing and Curse of Dimensionality in Safety Alignment
- PurpCode: Reasoning for Safer Code Generation
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
Related