RouteLLM: Learning to Route LLMs with Preference Data
2024/06/26 by Isaac Ong, Ong, Isaac, Amjad Almahairi +13 · 169 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2406.18665
openalex publication_date 2024/06/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) exhibit impressive capabilities across a wide range of tasks, yet the choice of which model to use often involves a trade-off between performance and cost. More powerful models, though effective, come with higher expenses, while less capable models are more cost-effective. To address this dilemma, we propose several efficient router models that dynamically select between a stronger and a weaker LLM during inference, aiming to optimize the balance between cost and response quality. We develop a training framework for these routers leveraging human preference data and data augmentation techniques to enhance performance. Our evaluation on widely-recognized benchmarks shows that our approach significantly reduces costs-by over 2 times in certain cases-without compromising the quality of responses. Interestingly, our router models also demonstrate significant transfer learning capabilities, maintaining their performance even when the strong and weak models are changed at test time. This highlights the potential of these routers to provide a cost-effective yet high-performance solution for deploying LLMs.
Cited by
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
- LLMBoost: Make Large Language Models Stronger with Boosting
- Explainable Model Routing for Agentic Workflows
- Kalypso: Relational LLM Serving
- WISERouter: LLM Routing with Workload Budget Constraint
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- Influence of Prompt Engineering on Small Language Models for Guarded Query Routing
- Codifying the Judge: Scalable Evaluation via Program Distillation
- Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
- CONCUR: A Framework for Continual Constrained and Unconstrained Routing
- Large Language Models as Generalist Policies for Network Optimization
- Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- Energy-Aware Data-Driven Model Selection in LLM-Orchestrated AI Systems
- Optimizing NetGPT via Routing-Based Synergy and Reinforcement Learning
- A note on conditional PAC-efficient reasoning in large language model routing
- PersonalizedRouter: Personalized LLM Routing via Graph-based User Preference Modeling
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
- CoLM: Collaborative Large Models via A Client-Server Paradigm
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
- Evaluation of Agents under Simulated AI Marketplace Dynamics
- The Stretto Execution Engine for LLM-Augmented Data Systems
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
- DiSRouter: Distributed Self-Routing for LLM Selections
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- Locket: Robust Feature-Locking Technique for Language Models
- Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm
- Don't Throw Away Your Pretrained Model
- NG-Router: Graph-Supervised Multi-Agent Collaboration for Nutrition Question Answering
- On the Provable Performance Guarantee of Efficient Reasoning Models
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- When to Reason: Semantic Router for vLLM
- Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs
- Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
- Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- Slm-mux: Orchestrating small language models for reasoning
- LLM Chemistry Estimation for Multi-LLM Recommendation
- Reward Model Routing in Alignment
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
- LLM Routing with Dueling Feedback
- Learning Compact Representations of LLM Abilities via Item Response Theory
- RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
- Reference-Free Rating of LLM Responses via Latent Information
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- LLM DNA: Tracing Model Evolution via Functional Representations
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- Fast Thinking for Large Language Models
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
- MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
- One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
- Latency and Token-Aware Test-Time Compute
- Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
- Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
- Delta Activations: A Representation for Finetuned Large Language Models
- Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
- Reasoning-Intensive Regression
- Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning
- Cost-Aware Contrastive Routing for LLMs
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
- Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
- Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
- LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
- RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
- Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
- Auto: The AGI Compiler
- LDP: An Identity-Aware Protocol for Multi-Agent LLM Systems
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
- Hybrid LLM Routing for Efficient App Feedback Classification
- Orchestration for Domain-specific Edge-Cloud Language Models
- Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
- Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration
- Economic Evaluation of LLMs
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- R1-Ranker: Teaching LLM Rankers to Reason
- Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
- Arch-Router: Aligning LLM Routing with Human Preferences
- Cost-Aware Routing for Efficient Text-To-Image Generation
- Keeping Up with the Models: Online Deployment and Routing of LLMs at Scale
- TagRouter: Learning Route to LLMs through Tags for Open-Domain Text Generation Tasks
- FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
- Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
- OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
- Answer Convergence as a Signal for Early Stopping in Reasoning
- Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- SkewRoute: Training-Free LLM Routing for Knowledge Graph Retrieval-Augmented Generation via Score Skewness of Retrieved Context
- Self-Route: Automatic Mode Switching via Capability Estimation for Efficient Reasoning
- Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
- VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- Guarded Query Routing for Large Language Models
- ThinkSwitcher: When to Think Hard, When to Think Fast
- Learnware of Language Models: Specialized Small Language Models Can Do Big
- Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers
- Thinkless: LLM Learns When to Think
- HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
- 402Pilot: An x402 Decision Layer for Autonomous Agent Micropayments
- Efficient RL Training for Reasoning Models via Length-Aware Optimization
- SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
- When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty
- Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- Internet of Agents: Fundamentals, Applications, and Challenges
- SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
- Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
- The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
- AI-Model Network: Concept, Current State and Future
- HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
- Separating Intelligence from Inference: A Standard for Edge-Native AI Computing
- When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
- Can a Crow Hatch a Falcon? Lineage Matters in Predicting Large Language Model Performance
- iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis
- Natural Language Query to Configuration for Retrieval Agents
- A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
- Credo: Declarative Control of LLM Pipelines via Beliefs and Policies
- RelayLLM: Efficient Reasoning via Collaborative Decoding
- History Matters: Meta-policy Delegation with Heterogeneous Multi-agent Reinforcement Learning
- COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
- Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination
- Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
- When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
- Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First
- Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation
- Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning
- Exploring How LLMs Capture and Represent Domain-Specific Knowledge
- Dynamic Early Exit in Reasoning Models
- EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
- Efficient Reasoning Models: A Survey
- EMAFusion: A Self-Optimizing System for Seamless LLM Selection and Integration
- Toward Super Agent System with Hybrid AI Routers
Related