vix.ing · top · new · best · stats · spec

BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation

2026/07/28 by Yinuo Zhu, He Liu, Boyuan Gu
#cs.HC

paper · pdf

Abstract

Large language models (LLMs) increasingly support biomedical research, yet their behavior under adversarial requests for dual-use knowledge remains insufficiently characterized. We introduce BioDisclose, a benchmark for measuring biomedical knowledge disclosure under adversarial elicitation. BioDisclose contains 480 prompts derived from 24 expert-authored scenarios across six biomedical risk domains and four elicitation families spanning academic, historical, role-playing, and decomposed prompting. We grade model responses on a four-level scale from refusal to executable disclosure, distinguishing high-level discussion from technically specific and actionable content, including refuse-then-leak behavior. Across five deployed LLM systems, detailed-or-higher disclosure rates vary substantially, ranging from 9.2% to 64.0%. Academic framing is the most effective elicitation family on average (43.2%), while laboratory safety scenarios show the highest disclosure rate across domains (51.5%). These results reveal pronounced variation across models, prompting strategies, and biomedical risk categories, suggesting that current safeguards remain uneven in high-stakes scientific settings. BioDisclose provides a focused testbed for evaluating biomedical safety beyond binary refusal metrics.

Citations

Related