WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
2024/06/26 by Seungju Han, Han, Seungju, Kavel Rao +13 · 182 citations
Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Safety Systems Engineering in Autonomy
paper · pdf · doi:10.48550/arxiv.2406.18495
openalex publication_date 2024/06/26 · openalex created_date 2024/06/28 · openalex updated_date 2026/07/28
Abstract
We introduce WildGuard -- an open, light-weight moderation tool for LLM safety that achieves three goals: (1) identifying malicious intent in user prompts, (2) detecting safety risks of model responses, and (3) determining model refusal rate. Together, WildGuard serves the increasing needs for automatic safety moderation and evaluation of LLM interactions, providing a one-stop tool with enhanced accuracy and broad coverage across 13 risk categories. While existing open moderation tools such as Llama-Guard2 score reasonably well in classifying straightforward model interactions, they lag far behind a prompted GPT-4, especially in identifying adversarial jailbreaks and in evaluating models' refusals, a key measure for evaluating safety behaviors in model responses. To address these challenges, we construct WildGuardMix, a large-scale and carefully balanced multi-task safety moderation dataset with 92K labeled examples that cover vanilla (direct) prompts and adversarial jailbreaks, paired with various refusal and compliance responses. WildGuardMix is a combination of WildGuardTrain, the training data of WildGuard, and WildGuardTest, a high-quality human-annotated moderation test set with 5K labeled items covering broad risk scenarios. Through extensive evaluations on WildGuardTest and ten existing public benchmarks, we show that WildGuard establishes state-of-the-art performance in open-source safety moderation across all the three tasks compared to ten strong existing open-source moderation models (e.g., up to 26.4% improvement on refusal detection). Importantly, WildGuard matches and sometimes exceeds GPT-4 performance (e.g., up to 3.9% improvement on prompt harmfulness identification). WildGuard serves as a highly effective safety moderator in an LLM interface, reducing the success rate of jailbreak attacks from 79.8% to 2.4%.
Cited by
- Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
- ProGuard: Towards Proactive Multimodal Safeguard
- Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
- AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models
- Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
- Semi-Supervised Learning for Large Language Models Safety and Content Moderation
- Safety Alignment of LMs via Non-cooperative Games
- AprielGuard
- Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
- CoPE: A Small Language Model for Steerable and Scalable Content Labeling
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
- Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
- From static to adaptive: immune memory-based jailbreak detection for large language models
- CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Benchmarking and Understanding Safety Risks in AI Character Platforms
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Fara-7B: An Efficient Agentic Model for Computer Use
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- SGuard-v1: Safety Guardrail for Large Language Models
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- Reimagining Safety Alignment with An Image
- Characterizing Selective Refusal Bias in Large Language Models
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
- Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
- Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
- Qwen3Guard Technical Report
- Budget-aware Test-time Scaling via Discriminative Verification
- Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- DeepResearchGuard: Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
- Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare
- Multimodal Policy Internalization for Conversational Agents
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Energy-Driven Steering: Reducing False Refusals in Large Language Models
- Auto-Prompt Ensemble for LLM Judge
- InvThink: Premortem Reasoning for Safer Language Models
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- NonTextual Target Attack
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Longitudinal Monitoring of LLM Content Moderation of Social Issues
- DSA, AIA, and LLMs: Approaches to conceptualizing and auditing moderation in LLM-based chatbots across languages and interfaces in the electoral contexts
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- A Comprehensive Evaluation framework of Alignment Techniques for LLMs
- Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
- In-Training Defenses against Emergent Misalignment in Language Models
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
- Libra: Large Chinese-based Safeguard for AI Content
- PurpCode: Reasoning for Safer Code Generation
- LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
- Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
- Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
- SAFER: Probing Safety in Reward Models with Sparse Autoencoder
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- Securing AI Systems: A Guide to Known Attacks and Impacts
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Aligning Spoken Dialogue Models from User Interactions
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- PL-Guard: Benchmarking Language Model Safety for Polish
- ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
- Improving Large Language Model Safety with Contrastive Representation Learning
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
- AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders
- Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
- OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
- PAM: Training Policy-Aligned Moderation Filters at Scale
- Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- Teaching Models to Understand (but not Generate) High-risk Data
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks
- Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
- Discovering Forbidden Topics in Language Models
- Refusal Direction is Universal Across Safety-Aligned Languages
- ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
- Sparse Activation Editing for Reliable Instruction Following in Narratives
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
- ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
- Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection
- Krikri: Advancing Open Large Language Models for Greek
- CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
- J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
- MergeBench: A Benchmark for Merging Domain-Specialized LLMs
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
- A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
- Practical Reasoning Interruption Attacks on Reasoning Large Language Models
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- Measuring and Eliminating Refusals in Military Large Language Models
- Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting
- Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
- PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
- Configurable Reward Model for Balanced Safety Alignment
- Trust The Typical
- Understanding Annotator Safety Policy with Interpretability
- Online Safety Monitoring for LLMs
- Do Thinking Tokens Help with Safety?
- What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
- There Is More to Refusal in Large Language Models than a Single Direction
- LLM Safety From Within: Detecting Harmful Content with Internal Representations
- Understanding and Mitigating Risks of Generative AI in Financial Services
- AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance
- ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems
- Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
- LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
- When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
- Safety in Large Reasoning Models: A Survey
- Social Pressure Breaks Majority Voting in LLM Safety Panels
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- PolicyGuard: Prompt-Configurable Semantic DLP for LLM Coding Agents
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
- aiXamine: Simplified LLM Safety and Security
- Detecting Safety Training Modification in Language Models via Activation Analysis
- Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
- The Structural Safety Generalization Problem
- Alleviating the Fear of Losing Alignment in LLM Fine-tuning
- X-Guard: Multilingual Guard Agent for Content Moderation
- Defense against Prompt Injection Attacks via Mixture of Encodings
- Geneshift: Impact of different scenario shift on Jailbreaking LLM
Related