vix.ing · top · new · best · stats

Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems

2025/09/30 by Aakriti Agrawal, Agrawal, Aakriti, Rohith Aralikatti +9 · 2 citations
Computer Science · #Arc (geometry) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Key (lock) #Machine Learning (cs.LG) #Matching (statistics) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Selection (genetic algorithm) #Task (project management) #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2510.02377

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/09/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Large Language Models (LLMs) have demonstrated exceptional capabilities, yet selecting the most reliable response from multiple LLMs remains a challenge, particularly in resource-constrained settings. Existing approaches often depend on costly external verifiers, human evaluators, or self-consistency techniques that require multiple samples from a single model. While multi-LLM systems produce more diverse responses than single models and thus have greater potential, they often underperform compared to single LLM self-consistency. We propose a principled, novel and computationally efficient method to select the best response from multiple different LLMs using a calibrated log-likelihood score, implicitly leveraging the inherent knowledge and confidence of these models. Our method demonstrates improvements of approx. 4%, 3%, and 5% across both debate (multi-round LLM discussions) and non-debate (Best-of-N with multiple LLMs) settings on GSM8K, MMLU (6 subsets), and ARC datasets respectively.

Citations

Cited by

Related