DeepInception: Hypnotize Large Language Model to Be Jailbreaker
2023/11/06 by Xuan Li, Li, Xuan, Zhanke Zhou +9 · 114 citations
Computer Science · #Adversarial Robustness in Machine Learning #Adversarial system #Artificial intelligence #Code (set theory) #Computer science #Computer security #Construct (python library) #Control (management) #Hate Speech and Cyberbullying Detection #Programming language #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2311.03191
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) have succeeded significantly in various applications but remain susceptible to adversarial jailbreaks that void their safety guardrails. Previous attempts to exploit these vulnerabilities often rely on high-cost computational extrapolations, which may not be practical or efficient. In this paper, inspired by the authority influence demonstrated in the Milgram experiment, we present a lightweight method to take advantage of the LLMs' personification capabilities to construct a virtual, nested scene, allowing it to realize an adaptive way to escape the usage control in a normal scenario. Empirically, the contents induced by our approach can achieve leading harmfulness rates with previous counterparts and realize a continuous jailbreak in subsequent interactions, which reveals the critical weakness of self-losing on both open-source and closed-source LLMs, e.g., Llama-2, Llama-3, GPT-3.5, GPT-4, and GPT-4o. The code and data are available at: https://github.com/tmlr-group/DeepInception.
Cited by
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- Do LLMs Know Their Vulnerable Scenarios?
- From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- Chain-of-Thought Hijacking
- Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Quantifying CBRN Risk in Frontier Models
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- BreakFun: Jailbreaking LLMs via Schema Exploitation
- FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands
- Imperceptible Jailbreaking against Large Language Models
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Security and Privacy Challenges of Large Language Models: A Survey
- Speculative Safety-Aware Decoding
- Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
- AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- LLM Jailbreak Detection for (Almost) Free!
- Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- What You Code Is What We Prove: Translating BLE App Logic into Formal Models with LLMs for Vulnerability Detection
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Mitigating Jailbreaks with Intent-Aware LLMs
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
- The Cost of Thinking: Increased Jailbreak Risk in Large Language Models
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
- Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety
- PhysicsEval: Inference-Time Techniques to Improve the Reasoning Proficiency of Large Language Models on Physics Problems
- MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems
- The bitter lesson of misuse detection
- MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
- TAI3: Testing Agent Integrity in Interpreting User Intent
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- HauntAttack: When Attack Follows Reasoning as a Shadow
- Exploring the Secondary Risks of Large Language Models
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- Should LLM Safety Be More Than Refusing Harmful Instructions?
- LLM Agents Should Employ Security Principles
- SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
- Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
- Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
- MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures -- A Comprehensive Framework
- Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
- Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Safety Alignment Can Be Not Superficial With Explicit Safety Signals
- Bullying the Machine: How Personas Increase LLM Vulnerability
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
- Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
- One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
- LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Attack and defense techniques in large language models: A survey and new perspectives
- Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
- Conflicts Make Large Reasoning Models Vulnerable to Attacks
- Ethical Risks in Deploying Large Language Models: An Evaluation of Medical Ethics Jailbreaking
- Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
- Safety in Large Reasoning Models: A Survey
- A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
- Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
- DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
- LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
Related