UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
2025/10/17 by Ning, Junzhi, Li, Wei, Tang, Cheng +24 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.15710
Abstract
Medical diagnostic applications require models that can process multimodal medical inputs (images, patient histories, lab results) and generate diverse outputs including both textual reports and visual content (annotations, segmentation masks, and images). Despite this need, existing medical AI systems disrupt this unified process: medical image understanding models interpret images but cannot generate visual outputs, while medical image generation models synthesize images but cannot provide textual explanations. This leads to gaps in data representation, feature integration, and task-level multimodal capabilities. To this end, we propose a multi-level framework that draws inspiration from diagnostic workflows through the Observation-Knowledge-Analysis (OKA) paradigm. Specifically, at the observation level, we construct UniMed-5M, a dataset comprising over 5.6M samples that reformat diverse unimodal data into multimodal pairs for foundational observation. At the knowledge level, we propose Progressive Curriculum Learning that systematically introduces medical multimodal knowledge. At the analysis level, we introduce UniMedVL, the first medical unified multimodal model for the simultaneous analysis of image understanding and generation tasks within a single architecture. UniMedVL achieves superior performance on five medical image understanding benchmarks, while matching specialized models in generation quality across eight medical imaging modalities. Crucially, our unified architecture enables bidirectional knowledge sharing: generation tasks enhance visual understanding features, demonstrating that integrating traditionally separate capabilities within a single medical framework unlocks improvements across diverse medical vision-language tasks. Code is available at https://github.com/uni-medical/UniMedVL.
Citations
- SAM 3: Segment Anything with Concepts
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- A Generative Foundation Model for Chest Radiography
- A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- MedGemma 1.5 Technical Report
- MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Emerging Properties in Unified Multimodal Pretraining
- RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
- Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
- VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
- GMAI-VL-R1: Harnessing Reinforcement Learning for Multimodal Medical Reasoning
- Towards Interpretable Counterfactual Generation via Multimodal Autoregression
- Unpaired Translation of Chest X-ray Images for Lung Opacity Diagnosis via Adaptive Activation Masks and Cross-Domain Alignment
- Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- SynthRAD2025 Grand Challenge dataset: generating synthetic CTs for radiotherapy
- HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
- GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI
- Anatomy-Guided Radiology Report Generation with Pathology-Aware Regional Prompts
- JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
- Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Deep Generative Models Unveil Patterns in Medical Images Through Vision-Language Conditioning
- Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
- ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
- Emu3: Next-Token Prediction is All You Need
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
- Multi-modal MRI Translation via Evidential Regression and Distribution Calibration
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
- All-In-One Medical Image Restoration via Task-Adaptive Routing
- CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- BraSyn 2023 challenge: Missing MRI synthesis and the effect of different learning objectives
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Pragmatic Radiology Report Generation
- Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects
- Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects
- NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding
- Improved Baselines with Visual Instruction Tuning
- MedSyn: Text-guided Anatomy-aware Synthesis of High-Fidelity 3D CT Images
- NExT-GPT: Any-to-Any Multimodal LLM
- PathLDM: Text conditioned Latent Diffusion Model for Histopathology
- Med-Flamingo: a Multimodal Medical Few-shot Learner
- Quilt-1M: One Million Image-Text Pairs for Histopathology
- XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- The Brain Tumor Segmentation (BraTS) Challenge 2023: Glioma Segmentation in Sub-Saharan Africa Patient Population (BraTS-Africa)
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical Documents
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- BCI: Breast Cancer Immunohistochemical Image Generation through Pyramid Pix2pix
- UVCGAN: UNet Vision Transformer cycle-consistent GAN for unpaired image-to-image translation
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- SwinIR: Image Restoration Using Swin Transformer
- SwinIR: Image Restoration Using Swin Transformer
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- MedICaT: A Dataset of Medical Images, Captions, and Textual References
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- Diverse Image-to-Image Translation via Disentangled Representations
- Multimodal Unsupervised Image-to-Image Translation
- High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs
- Unsupervised Image-to-Image Translation Networks
- Image-to-Image Translation with Conditional Adversarial Networks
- Accurate Image Super-Resolution Using Very Deep Convolutional Networks
- Image Super-Resolution Using Deep Convolutional Networks
- Image Super-Resolution Using Deep Convolutional Networks
Cited by
Related