vix.ing · top · new · best · stats · spec

EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery

2025/09/19 by Yang Wennuo, Wang, Gui, Xudong Ma +16
Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Retinal Imaging and Analysis

paper · pdf · doi:10.48550/arxiv.2509.15596

openalex publication_date 2025/09/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap, we develop EyePCR, a large-scale benchmark for ophthalmic surgery analysis, grounded in structured clinical knowledge to evaluate cognition across Perception, Comprehension and Reasoning. EyePCR offers a richly annotated corpus with more than 210k VQAs, which cover 1048 fine-grained attributes for multi-view perception, medical knowledge graph of more than 25k triplets for comprehension, and four clinically grounded reasoning tasks. The rich annotations facilitate in-depth cognitive analysis, simulating how surgeons perceive visual cues and combine them with domain knowledge to make decisions, thus greatly improving models' cognitive ability. In particular, EyePCR-MLLM, a domain-adapted variant of Qwen2.5-VL-7B, achieves the highest accuracy on MCQs for Perception among compared models and outperforms open-source models in Comprehension and Reasoning, rivalling commercial models like GPT-4.1. EyePCR reveals the limitations of existing MLLMs in surgical cognition and lays the foundation for benchmarking and enhancing clinical reliability of surgical video understanding models.

Citations

Related