SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
2025/10/17 by Hong, Hanbin, Feng, Shuya, Naderloui, Nima +6 · 1 citation
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.15476
Abstract
Large Language Models (LLMs) have rapidly become integral to real-world applications, powering services across diverse sectors. However, their widespread deployment has exposed critical security risks, particularly through jailbreak prompts that can bypass model alignment and induce harmful outputs. Despite intense research into both attack and defense techniques, the field remains fragmented: definitions, threat models, and evaluation criteria vary widely, impeding systematic progress and fair comparison. In this Systematization of Knowledge (SoK), we address these challenges by (1) proposing a holistic, multi-level taxonomy that organizes attacks, defenses, and vulnerabilities in LLM prompt security; (2) formalizing threat models and cost assumptions into machine-readable profiles for reproducible evaluation; (3) introducing an open-source evaluation toolkit for standardized, auditable comparison of attacks and defenses; (4) releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date;\footnoteThe dataset is released at \hrefhttps://huggingface.co/datasets/youbin2014/JailbreakDB\textcolorpurplehttps://huggingface.co/datasets/youbin2014/JailbreakDB. and (5) presenting a comprehensive evaluation platform and leaderboard of state-of-the-art methods \footnotewill be released soon.. Our work unifies fragmented research, provides rigorous foundations for future studies, and supports the development of robust, trustworthy LLMs suitable for high-stakes deployment.
Citations
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
- Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
- SoK: Prompt Hacking of Large Language Models
- BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
- RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
- JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
- Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
- Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
- Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
- FlipAttack: Jailbreak LLMs via Flipping
- PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
- PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
- AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
- Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection
- HSF: Defending against Jailbreak Attacks with Hidden State Filtering
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
- Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
- Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
- A Jailbroken GenAI Model Can Cause Substantial Harm: GenAI-powered Applications are Vulnerable to PromptWares
- Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
- LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
- PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
- RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
- Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
- Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
- Does Refusal Training in LLMs Generalize to the Past Tense?
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
- Automated Progressive Red Teaming
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
- Badllama 3: removing safety finetuning from Llama 3 in minutes
- Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
- Poisoned LangChain: Jailbreak LLMs by LangChain
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
- Adversaries Can Misuse Combinations of Safe Models
- garak: A Framework for Security Probing Large Language Models
- RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
- JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
- Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
- Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
- SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
- Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
- Improving Alignment and Robustness with Circuit Breakers
- Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
- Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
- Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
- Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
- Improved Generation of Adversarial Examples Against Safety-aligned LLMs
- Are PPO-ed Language Models Hackable?
- Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
- Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
- GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
- Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
- CLARE: Cognitive Load Assessment in REaltime with Multimodal Data
- Don't Say No: Jailbreaking LLM by Suppressing Refusal
- MoDE: CLIP Data Experts via Clustering
- Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Optimization-based Prompt Injection Attack to LLM-as-a-Judge
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
- Detoxifying Large Language Models via Knowledge Editing
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
- EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
- AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
- Distract Large Language Models for Automatic Jailbreak Attack
- CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
- Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
- Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- Defending LLMs against Jailbreaking Attacks via Backtranslation
- DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
- PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
- How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
- Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
- Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
- Coercing LLMs to do and reveal (almost) anything
- A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
- Defending Jailbreak Prompts via In-Context Adversarial Game
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
- Query-Based Adversarial Prompt Generation
- SPML: A DSL for Defending Language Models Against Prompt Attacks
- PAL: Proxy-Guided Black-Box Attack on Large Language Models
- Attacking Large Language Models with Projected Gradient Descent
- Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
- SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
- Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
- Comprehensive Assessment of Jailbreak Attacks Against LLMs
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
- Weak-to-Strong Jailbreaking on Large Language Models
- A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
- PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models
- Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Evil Geniuses: Delving into the Safety of LLM-based Agents
- Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
- A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
- Attack Prompt Generation for Red Teaming and Defending Large Language Models
- Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- Multilingual Jailbreak Challenges in Large Language Models
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Low-Resource Languages Jailbreak GPT-4
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Goal-Oriented Prompt Attack and Safety Evaluation for LLMs
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- RAIN: Your Language Models Can Align Themselves without Finetuning
- Certifying LLM Safety against Adversarial Prompting
- Open Sesame! Universal Black Box Jailbreaking of Large Language Models
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Detecting Language Model Attacks with Perplexity
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Platypus: Quick, Cheap, and Powerful Refinement of LLMs
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
- BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- Jailbroken: How Does LLM Safety Training Fail?
- From ChatGPT to ThreatGPT: Impact of Generative AI in Cybersecurity and Privacy
- Prompt Injection attack against LLM-integrated Applications
- Adversarial Demonstration Attacks on Large Language Models
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- SneakyPrompt: Jailbreaking Text-to-image Generative Models
- Multi-step Jailbreaking Privacy Attacks on ChatGPT
- Instruction Tuning with GPT-4
- Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
- Privacy of federated QR decomposition using additive secure multiparty computation
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Gradient-based Adversarial Attacks against Text Transformers
- Measuring Massive Multitask Language Understanding
- Universal Adversarial Triggers for Attacking and Analyzing NLP
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
- Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
- Intention Analysis Makes LLMs A Good Jailbreak Defender
- On Prompt-Driven Safeguarding for Large Language Models
- Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
- JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
Cited by
Related