vix.ing · top · new · best · stats · spec

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

2024/06/07 by Lukas Helff, F. Friedrich, Helff, Lukas +7 · 23 citations
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Object Detection Techniques #Machine Learning (cs.LG) #Retinal Imaging and Analysis

paper · pdf · doi:10.48550/arxiv.2406.05113

openalex publication_date 2024/06/07 · openalex created_date 2024/06/11 · openalex updated_date 2026/07/28

Abstract

This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this end, we establish a novel open framework, describing a customizable safety taxonomy, data preprocessing, augmentation, and training setup. For teaching a VLM safeguard on safety, we further create a multimodal safety dataset with high-quality human expert annotations, where each image is labeled with a safety rating, category, and rationale. We also employ advanced augmentations to support context-specific assessments. The resulting LlavaGuard models, ranging from 0.5B to 7B, serve as a versatile tool for evaluating the safety compliance of visual content against flexible policies. In comprehensive experiments, LlavaGuard outperforms both state-of-the-art safeguards and VLMs in accuracy and in flexibly handling different policies. Additionally, we demonstrate LlavaGuard's performance in two real-world applications: large-scale dataset annotation and moderation of text-to-image models. We make our entire framework, including the dataset, model weights, and training code.

Cited by

Related