MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
2025/05/17 by Jingkun Yue, Yue, Jingkun, Siqi Zhang +11 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2505.11852
openalex publication_date 2025/05/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domains. While existing medical visual grounding benchmarks primarily focus on single-image scenarios, real-world clinical applications often involve sequential images, where accurate lesion localization across different modalities and temporal tracking of disease progression (e.g., pre- vs. post-treatment comparison) require fine-grained cross-image semantic alignment and context-aware reasoning. To remedy the underrepresentation of image sequences in existing medical visual grounding benchmarks, we propose MedSG-Bench, the first benchmark tailored for Medical Image Sequences Grounding. It comprises eight VQA-style tasks, formulated into two paradigms of the grounding tasks, including 1) Image Difference Grounding, which focuses on detecting change regions across images, and 2) Image Consistency Grounding, which emphasizes detection of consistent or shared semantics across sequential images. MedSG-Bench covers 76 public datasets, 10 medical imaging modalities, and a wide spectrum of anatomical structures and diseases, totaling 9,630 question-answer pairs. We benchmark both general-purpose MLLMs (e.g., Qwen2.5-VL) and medical-domain specialized MLLMs (e.g., HuatuoGPT-vision), observing that even the advanced models exhibit substantial limitations in medical sequential grounding tasks. To advance this field, we construct MedSG-188K, a large-scale instruction-tuning dataset tailored for sequential visual grounding, and further develop MedSeq-Grounder, an MLLM designed to facilitate future research on fine-grained understanding across medical sequential images. The benchmark, dataset, and model are available at https://huggingface.co/MedSG-Bench
Citations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MedGemma Technical Report
- DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- FreeTumor: Large-Scale Generative Tumor Synthesis in Computed Tomography Images for Improving Tumor Recognition
- Qwen2.5-VL Technical Report
- MMXU: A Multi-Modal and Multi-X-ray Understanding Dataset for Disease Progression
- Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
- Toward Visual Grounding: A Survey
- Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
- BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
- ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis
- GPT-4o System Card
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
- LLaVA-OneVision: Easy Visual Task Transfer
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- MAIRA-2: Grounded Radiology Report Generation
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding
- PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- QUBIQ: Uncertainty Quantification for Biomedical Image Segmentation Challenge
- Towards Generalizable Tumor Synthesis
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- SegRap2023: A Benchmark of Organs-at-Risk and Gross Tumor Volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
- GSVA: Generalized Segmentation via Multimodal Large Language Models
- PixelLM: Pixel Reasoning with Large Multimodal Model
- NExT-Chat: An LMM for Chat, Detection and Segmentation
- GLaMM: Pixel Grounding Large Multimodal Model
- LISA: Reasoning Segmentation via Large Language Model
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Efficient automatic segmentation for multi-level pulmonary arteries: The PARSE challenge
- Micrograph segmentations for DDEVD
- Segment Anything
- Label-Free Liver Tumor Segmentation
- GPT-4 Technical Report
- Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- BayeSeg: Bayesian Modeling for Medical Image Segmentation with Interpretable Generalizability
- The state-of-the-art 3D anisotropic intracranial hemorrhage segmentation on non-contrast head CT: The INSTANCE challenge
- The Extreme Cardiac MRI Analysis Challenge under Respiratory Motion (CMRxMotion)
- AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Surgical Automation in Laparoscopic Hysterectomy
- CTooth+: A Large-scale Dental Cone Beam Computed Tomography Dataset and Benchmark for Tooth Volume Segmentation
- CTooth: A Fully Annotated 3D Dataset and Benchmark for Tooth Volume Segmentation on Cone Beam Computed Tomography Images
- AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation
- VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography
- ImageTBAD: A 3D Computed Tomography Angiography Image Dataset for Automatic Segmentation of Type-B Aortic Dissection
- Deep Learning methods for automatic evaluation of delayed enhancement-MRI. The results of the EMIDEC challenge
- CTSpine1K: A Large-Scale Dataset for Spinal Vertebrae Segmentation in Computed Tomography
- CutPaste: Self-Supervised Learning for Anomaly Detection and Localization
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- Deep Learning to Segment Pelvic Bones: Large-scale CT Datasets and Baseline Models
- AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem?
- Kvasir-Instrument: Diagnostic and therapeutic tool segmentation dataset in gastrointestinal endoscopy
- Convolutional Sparse Support Estimator Based Covid-19 Recognition from\n X-ray Images
- AGE Challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
- Computer Aided Detection for Pulmonary Embolism Challenge (CAD-PE)
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)
- Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge
- Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved?
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)
- GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
- Development of a Digital Image Database for Chest Radiographs With and Without a Lung Nodule
Cited by
Related