LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
2025/09/30 by Huang, Guolei, Qinzhi Peng, Peng, Qinzhi +7
Psychology · Computer Science · #Language, Metaphor, and Cognition #Speech and dialogue systems #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2509.25896
Abstract
As Vision-Language Models (VLMs) move into interactive, multi-turn use, safety concerns intensify for multimodal multi-turn dialogue, which is characterized by concealment of malicious intent, contextual risk accumulation, and cross-modal joint risk. These characteristics limit the effectiveness of content moderation approaches designed for single-turn or single-modality settings. To address these limitations, we first construct the Multimodal Multi-turn Dialogue Safety (MMDS) dataset, comprising 4,484 annotated dialogues and a comprehensive risk taxonomy with 8 primary and 60 subdimensions. As part of MMDS construction, we introduce Multimodal Multi-turn Red Teaming (MMRT), an automated framework for generating unsafe multimodal multi-turn dialogues. We further propose LLaVAShield, which audits the safety of both user inputs and assistant responses under specified policy dimensions in multimodal multi-turn dialogues. Extensive experiments show that LLaVAShield significantly outperforms state-of-the-art VLMs and existing content moderation tools while demonstrating strong generalization and flexible policy adaptation. Additionally, we analyze vulnerabilities of mainstream VLMs to harmful inputs and evaluate the contribution of key components, advancing understanding of safety mechanisms in multimodal multi-turn dialogues.
Citations
- OpenAI GPT-5 System Card
- Qwen3-VL Technical Report
- Qwen3Guard Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Multi-Turn Jailbreaks Are Simpler Than They Seem
- Many-Turn Jailbreaking
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
- SafeCoT: Improving VLM Safety with Minimal Reasoning
- SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
- Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
- VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
- ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- Qwen3 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- ShieldGemma 2: Robust and Tractable Image Content Moderation
- MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning
- Survey of Adversarial Robustness in Multimodal Large Language Models
- BingoGuard: LLM Content Moderation Tools with Risk Levels
- Qwen2.5-VL Technical Report
- Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
- Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- GuardReasoner: Towards Reasoning-based LLM Safeguards
- RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
- Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
- GPT-4o System Card
- SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
- Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
- Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles
- LLaVA-OneVision: Easy Visual Task Transfer
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
- MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
- MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
- LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- Safety Alignment for Vision Language Models
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
- CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- On the Robustness of Large Multimodal Models Against Image Adversarial Attacks
- MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Large language models in medicine
- ChatGPT for good? On opportunities and challenges of large language models for education
- Talking About Large Language Models
- Talking about Large Language Models
- A Holistic Approach to Undesired Content Detection in the Real World
- A New Generation of Perspective API: Efficient Multilingual Character-level Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- Learning from the Worst: Dynamically Generated Datasets to Improve\n Online Hate Detection
- RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
Related