Jailbreaking Black Box Large Language Models in Twenty Queries
2023/10/12 by Patrick Chao, Chao, Patrick, Alexander Robey +9 · 1 voice · 423 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.08419
openalex publication_date 2023/10/12 · arxiv published 2023/10/12 · arxiv updated 2024/07/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
There is growing interest in ensuring that large language models (LLMs) align with human values. However, the alignment of such models is vulnerable to adversarial jailbreaks, which coax LLMs into overriding their safety guardrails. The identification of these vulnerabilities is therefore instrumental in understanding inherent weaknesses and preventing future misuse. To this end, we propose Prompt Automatic Iterative Refinement (PAIR), an algorithm that generates semantic jailbreaks with only black-box access to an LLM. PAIR -- which is inspired by social engineering attacks -- uses an attacker LLM to automatically generate jailbreaks for a separate targeted LLM without human intervention. In this way, the attacker LLM iteratively queries the target LLM to update and refine a candidate jailbreak. Empirically, PAIR often requires fewer than twenty queries to produce a jailbreak, which is orders of magnitude more efficient than existing algorithms. PAIR also achieves competitive jailbreaking success rates and transferability on open and closed-source LLMs, including GPT-3.5/4, Vicuna, and Gemini.
Cited by
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- Boundary Point Jailbreaking of Black-Box LLMs
- Do LLMs Know Their Vulnerable Scenarios?
- TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- AgentWatcher: A Rule-based Prompt Injection Monitor
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- Safety Alignment of LMs via Non-cooperative Games
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
- Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
- Are Your Agents Upward Deceivers?
- In-Context Representation Hijacking
- From static to adaptive: immune memory-based jailbreak detection for large language models
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- An Invariant Latent Space Perspective on Language Model Inversion
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- Robustness of LLM-enabled vehicle trajectory prediction under data security threats
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- Efficient LLM Safety Evaluation through Multi-Agent Debate
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
- Jailbreaking in the Haystack
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
- Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
- Reimagining Safety Alignment with An Image
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- Diffusion LLMs are Natural Adversaries for any LLM
- Reasoning Up the Instruction Ladder for Controllable Language Models
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- A Survey on Unlearning in Large Language Models
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
- Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
- Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- Black-box Optimization of LLM Outputs by Asking for Directions
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
- On the Convergence of Moral Self-Correction in Large Language Models
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
- Learning to Interpret Weight Differences in Language Models
- Imperceptible Jailbreaking against Large Language Models
- Proactive defense against LLM Jailbreak
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
- NonTextual Target Attack
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- CHAI: Command Hijacking against embodied AI
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
- STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
- SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
- Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- Compliance2LoRA: Personalizable On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- AI Agents May Always Fall for Prompt Injections
- Security and Privacy Challenges of Large Language Models: A Survey
- Adversarial machine learning :
- Speculative Safety-Aware Decoding
- Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- Semantic Representation Attack against Aligned Large Language Models
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- MillStone: How Open-Minded Are LLMs?
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- Evaluating Language Model Reasoning about Confidential Information
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
- Mitigating Jailbreaks with Intent-Aware LLMs
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Many-Turn Jailbreaking
- Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks
- JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
- Automatic LLM Red Teaming
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
- Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- Activation-Guided Local Editing for Jailbreaking Attacks
- Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
- Strategic Deflection: Defending LLMs from Logit Manipulation
- Enhancing Jailbreak Attacks on LLMs via Persona Prompts
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
- A Dynamic Stackelberg Game Framework for Agentic AI Defense Against LLM Jailbreaking
- GPUHammer: Rowhammer Attacks on GPU Memories are Practical
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
- GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- A Mathematical Theory of Discursive Networks
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
- CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
- Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
- The bitter lesson of misuse detection
- Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
- The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
- Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
- Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion
- Who's the Mole? Modeling and Detecting Intention-Hiding Malicious Agents in LLM-Based Multi-Agent Systems
- Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
- Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
- Large Language Model-Driven Closed-Loop UAV Operation with Semantic Observations
- Reasoning as an Adaptive Defense for Safety
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
- A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
- Securing AI Systems: A Guide to Known Attacks and Impacts
- On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
- VERA: Variational Inference Framework for Jailbreaking Large Language Models
- MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
- Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- TAI3: Testing Agent Integrity in Interpreting User Intent
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
- Command-V: Pasting LLM Behaviors via Activation Profiles
- GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
- Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- Probing the Robustness of Large Language Models Safety to Latent Perturbations
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- LLM Jailbreak Oracle
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
- RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
- Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
- Mind the Web: The Security of Web Use Agents
- Jailbreak Transferability Emerges from Shared Representations
- Universal Jailbreak Suffixes Are Strong Attention Hijackers
- SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
- Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
- Building Trustworthy AI by Addressing its 16+2 Desiderata with Goal-Directed Commonsense Reasoning
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- Exploring the Secondary Risks of Large Language Models
- Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
- Improving Large Language Model Safety with Contrastive Representation Learning
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- Benchmarking Misuse Mitigation Against Covert Adversaries
- RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- Adversarial Attacks on Robotic Vision Language Action Models
- Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- The Security Threat of Compressed Projectors in Large Vision-Language Models
- From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
- COSMIC: Generalized Refusal Direction Identification in LLM Activations
- A Red Teaming Roadmap Towards System-Level Safety
- LLM Agents Should Employ Security Principles
- Position: Federated Foundation Language Model Post-Training Should Focus on Open-Source Models
- MEF: A Capability-Aware Multi-Encryption Framework for Evaluating Vulnerabilities in Black-Box Large Language Models
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
- SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA
- GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- Jailbreak Distillation: Renewable Safety Benchmarking
- Concealment of Intent: A Game-Theoretic Analysis
- Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
- Improved Representation Steering for Language Models
- PAM: Training Policy-Aligned Moderation Filters at Scale
- When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
- Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
- Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
- Lifelong Safety Alignment for Language Models
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters
- Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
- Security Concerns for Large Language Models: A Survey
- Mitigating Deceptive Alignment via Self-Monitoring
- Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
- Safety Alignment via Constrained Knowledge Unlearning
- PD3F: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
- Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
- Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
- Emergent Standing Wave Dynamics and Attractor Basins in Transformer Latent Spaces via Prompt Driven ConstraintsAnathema to Corporate Control by Ingrid Johnson
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
- CAIN: Hijacking LLM-Humans Conversations via Malicious System Prompts
- MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
- Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
- How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
- Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
- Causes and Consequences of Representational Similarity in Machine Learning Models
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
- SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
- Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
- Soft Prompts for Evaluation: Measuring Conditional Distance of Capabilities
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
- Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
- Bullying the Machine: How Personas Increase LLM Vulnerability
- A Survey of Attacks on Large Language Models
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
- JULI: Jailbreak Large Language Models by Self-Introspection
- Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
- Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
- Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
- One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
- Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- Attack and defense techniques in large language models: A survey and new perspectives
- Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
- Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
- Assessing Automated Prompt Injection Attacks in Agentic Environments
- Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- OET: Optimization-based prompt injection Evaluation Toolkit
- Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement
- Prompt Injection as Role Confusion
- A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models
- A New Framework for Cybersecurity Refusals in AI Agents
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
- Trust The Typical
- XBreaking: Understanding how LLMs security alignment can be broken
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- ACE: A Security Architecture for LLM-Integrated App Systems
- Estimating Tail Risks in Language Model Output Distributions
- NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
- JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
- The Automation Advantage in AI Red Teaming
- Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
- RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
- Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
- Conflicts Make Large Reasoning Models Vulnerable to Attacks
- PIArena: A Platform for Prompt Injection Evaluation
- Ethical Risks in Deploying Large Language Models: An Evaluation of Medical Ethics Jailbreaking
- ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
- Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
- Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
- Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
- Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
- Safety Pretraining: Toward the Next Generation of Safe AI
- A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
- Inducing Vulnerable Code Generation in LLM Coding Assistants
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
- RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
- DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
- Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
- AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
- GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
- Antidistillation Sampling
- Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks
- Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
- RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
- SaRO: Enhancing LLM Safety through Reasoning-based Alignment
- Feature-Aware Malicious Output Detection and Mitigation
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- Geneshift: Impact of different scenario shift on Jailbreaking LLM
- Bypassing Safety Guardrails in LLMs Using Humor
- Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
Discussions
Related