Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
2025/06/09 by Ruiyang Zhang, Zhang, Ruiyang, Hu Zhang +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2506.07575
openalex publication_date 2025/06/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Multimodal Models (LMMs), harnessing the complementarity among diverse modalities, are often considered more robust than pure Language Large Models (LLMs); yet do LMMs know what they do not know? There are three key open questions remaining: (1) how to evaluate the uncertainty of diverse LMMs in a unified manner, (2) how to prompt LMMs to show its uncertainty, and (3) how to quantify uncertainty for downstream tasks. In an attempt to address these challenges, we introduce Uncertainty-o: (1) a model-agnostic framework designed to reveal uncertainty in LMMs regardless of their modalities, architectures, or capabilities, (2) an empirical exploration of multimodal prompt perturbations to uncover LMM uncertainty, offering insights and findings, and (3) derive the formulation of multimodal semantic uncertainty, which enables quantifying uncertainty from multimodal responses. Experiments across 18 benchmarks spanning various modalities and 10 LMMs (both open- and closed-source) demonstrate the effectiveness of Uncertainty-o in reliably estimating LMM uncertainty, thereby enhancing downstream tasks such as hallucination detection, hallucination mitigation, and uncertainty-aware Chain-of-Thought reasoning.
Citations
- Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
- Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
- I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
- A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions
- VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
- GPT-4o System Card
- PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
- Visual Prompting in Multimodal Large Language Models: A Survey
- A Survey on Evaluation of Multimodal Large Language Models
- Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection
- RGB2Point: 3D Point Cloud Generation from Single RGB Images
- Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
- Uncertainty Aware Learning for Language Model Alignment
- Can Graph Learning Improve Planning in LLM-based Agents?
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
- Hallucination of Multimodal Large Language Models: A Survey
- Enhancing Video-Language Representations With Structural Spatio-Temporal Alignment
- MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
- A Review of Multi-Modal Large Language and Vision Models
- UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
- Visual Hallucinations of Multi-modal Large Language Models
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- The Revolution of Multimodal Large Language Models: A Survey
- Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
- Unified Hallucination Detection for Multimodal Large Language Models
- UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
- MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Prompt Highlighter: Interactive Control for Multi-Modal LLMs
- OneLLM: One Framework to Align All Modalities with Language
- CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
- Self-correcting LLM-controlled Diffusion Models
- Multimodal Large Language Models: A Survey
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- NExT-GPT: Any-to-Any Multimodal LLM
- PointLLM: Empowering Large Language Models to Understand Point Clouds
- Evaluation and Analysis of Hallucination in Large Vision-Language Models
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Multimodal Automated Fact-Checking: A Survey
- Any-to-Any Generation via Composable Diffusion
- Evaluating Object Hallucination in Large Vision-Language Models
- ImageBind: One Embedding Space To Bind Them All
- Visual Instruction Tuning
- Human Uncertainty in Concept-Based AI Systems
- GPT-4 Technical Report
- VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation
- Point-E: A System for Generating 3D Point Clouds from Complex Prompts
- Objaverse: A Universe of Annotated 3D Objects
- Robust Speech Recognition via Large-Scale Weak Supervision
- MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
- AudioGen: Textually Guided Audio Generation
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- Denoising Diffusion Probabilistic Models
- Language Models are Few-Shot Learners
- Evaluating the Evaluation of Diversity in Natural Language Generation
- Clotho: An Audio Captioning Dataset
- Human uncertainty makes classification more robust
- Object Hallucination in Image Captioning
- Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling
- ShapeNet: An Information-Rich 3D Model Repository
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for\n Richer Image-to-Sentence Models
- Microsoft COCO Captions: Data Collection and Evaluation Server
Related