2025/04/01 by Jianhao Chen, Chen, Jianhao, Bocheng Zhou +18 · 12 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Consistency (knowledge bases) #FOS: Computer and information sciences #Inference #Leverage (statistics) #Machine Learning in Healthcare #Sampling (signal processing) #Scaling #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2504.00762
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
This paper presents a simple, effective, and cost-efficient strategy to improve LLM performance by scaling test-time compute. Our strategy builds upon the repeated-sampling-then-voting framework, with a novel twist: incorporating multiple models, even weaker ones, to leverage their complementary strengths that potentially arise from diverse training data and paradigms. By using consistency as a signal, our strategy dynamically switches between models. Theoretical analysis highlights the efficiency and performance advantages of our strategy. Extensive experiments on six datasets demonstrate that our strategy not only outperforms self-consistency and state-of-the-art multi-agent debate approaches, but also significantly reduces inference costs. Additionally, ModelSwitch requires only a few comparable LLMs to achieve optimal performance and can be extended with verification methods, demonstrating the potential of leveraging multiple LLMs in the generation-verification paradigm.