vix.ing · top · new · best · stats · spec

Deciphering the “non-verbal” code: A preliminary exploration of multimodal large language models for neonatal pain recognition

2026/02/01 by Xiaosong Jiang, Nannan Yang, Tingqi Shi +4 · 1 voice
Medicine · #Pediatric Pain Management Techniques #Pain Management and Opioid Use #Neonatal and fetal brain pathology

paper · doi:10.1177/20552076261473729

openalex publication_date 2026/02/01 · openalex created_date 2026/07/25 · openalex updated_date 2026/07/28

Abstract

Background: Neonatal pain assessment primarily relies on behavioral rating scales. Although these tools provide objective measures, their scores are susceptible to inter-rater variability, potentially introducing bias. Moreover, intermittent assessments cannot capture pain progression over time. Multimodal large language models (MLLMs) offer a novel approach to addressing these limitations; however, their effectiveness in neonatal pain recognition remains largely unexplored. Objective: This study aimed to evaluate the performance of several MLLMs in neonatal pain video classification and investigate their applicability, limitations, and potential for clinical implementation. Methods: A previously established, annotation-validated multimodal dataset of acute neonatal pain was used. The dataset comprised 426 video recordings of neonatal heel lance procedures. Three MLLMs with dynamic video analysis capabilities (Qwen3-VL-Plus, Gemini-3-Pro, and ERNIE-4.5-Turbo) were evaluated. Carefully designed prompts guided the models to simulate nurses' pain assessments. Model performance was comprehensively evaluated using accuracy, weighted Cohen's kappa coefficient, precision, recall, and F1 score. Results: Gemini-3-Pro achieved the best classification performance, with an accuracy of 86.9% and a weighted kappa coefficient of 0.695, followed by Qwen3-VL-Plus (83.6%, kappa = 0.624) and ERNIE-4.5-Turbo (78.4%, kappa = 0.563). Gemini-3-Pro demonstrated an excellent recall of 0.974 and an F1 score of 0.905 for the pain category. However, all models showed relatively lower recall for the no-pain category (0.664-0.868), suggesting a trade-off between pain and no-pain classification and indicating challenges in achieving both high sensitivity and specificity. Conclusion: MLLMs demonstrate considerable potential for neonatal pain assessment, with particularly strong performance in pain screening. Their structured reasoning and interpretable outputs may support clinical decision-making. However, the observed performance imbalance between pain and no-pain classification suggests that these models are not yet capable of independently replacing clinical judgment and are better suited as assistive tools to optimize clinical workflows.

Citations

Discussions