2026/06/25 by Ting Fang Tan, Kabilan Elangovan, JASMINE ONG +12 · 1 voice
Medicine · Health Professions · Computer Science · #Artificial Intelligence in Healthcare and Education #Electronic Health Records Systems #Machine Learning in Healthcare #Consistency (knowledge bases) #Documentation #Reliability (semiconductor) #Health care #Domain (mathematical analysis) #Subject-matter expert #Quality (philosophy) #Internal consistency
paper · doi:10.1016/j.xcrm.2026.102883
openalex publication_date 2026/06/25 · openalex created_date 2026/06/26 · openalex updated_date 2026/07/29
Domain-specific evaluation is essential for clinical validation. We propose S.C.O.R.E. (Safety, Consensus & Context, Objectivity, Reproducibility, Explainability), a five-dimensional framework for structured expert evaluation of LLM-generated healthcare responses. S.C.O.R.E. has been validated against quantitative metrics (BLEU, ROUGE, and BERTScore) using three LLMs (GPT-4o, Claude 4 Sonnet, and DeepSeek) across ophthalmology, medication, and anesthesia. While quantitative metrics frequently misclassified clinically appropriate responses as inaccurate, S.C.O.R.E. demonstrated acceptable internal consistency (Cronbach's α 0.745) in the hyperparameter-optimized domain and detected large effect sizes (Cliff's δ 0.68-0.92) reflecting optimization status. Model rankings reversed across specialties-GPT-4o excelled in ophthalmology (optimized domain), while others in non-optimized domains-thus domain-specific tuning is both necessary and detectable through expert evaluation. S.C.O.R.E.'s correlation between framework reliability and optimization status validates its utility for iterative model refinement. This structured approach enables practical clinical validation, providing actionable feedback for developers and supporting regulatory compliance through standardized documentation of safety, evidence alignment, and explainability.