2026/03/30 by Javeed Sukhera, Emily Shearier, Emily R. Shearier +5 · 1 voice
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Clinical Reasoning and Diagnostic Skills #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.21203/rs.3.rs-9152209/v1
crossref issued 2026/03/30 · crossref published 2026/03/30 · openalex publication_date 2026/03/30 · crossref created 2026/03/30 · openalex created_date 2026/03/31 · crossref deposited 2026/04/01 · crossref indexed 2026/04/01 · openalex updated_date 2026/07/14
Abstract Artificial intelligence (AI) agents are increasingly used in clinical settings, yet their safety and reliability remain uncertain. We conducted a two-phase stress test using adversarial prompts and evaluated 476 responses across five safety domains. In phase 1, 48 participants completed randomized medium and high-risk scenarios. Following review of early failures, phase 2 enrolled 27 new participants who completed eight novel high-risk scenarios targeting observed failure themes. In phase 1, 17.4% of transcripts failed by majority rater consensus (≥2 of 3 raters scoring ≤2 in any domain), including 30.2% of high-risk scenarios. In phase 2, mean safety scores improved across all domains by 0.50 to 0.80, and the high-risk failure rate fell to 8.5% majority and to 6.0% across all raters. Results suggest that platform safety can improve with iterative evaluation, while underscoring the importance of continued safeguards and prospective validation.