A Survey on Collaborative Mechanisms Between Large and Small Language Models
2025/05/12 by Yi Chen, Chen, Yi, JiaHao Zhao +3 · 6 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Big Data and Digital Economy #Computation and Language (cs.CL) #FOS: Computer and information sciences #IoT and Edge/Fog Computing
paper · pdf · doi:10.48550/arxiv.2505.07460
openalex publication_date 2025/05/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) deliver powerful AI capabilities but face deployment challenges due to high resource costs and latency, whereas Small Language Models (SLMs) offer efficiency and deployability at the cost of reduced performance. Collaboration between LLMs and SLMs emerges as a crucial paradigm to synergistically balance these trade-offs, enabling advanced AI applications, especially on resource-constrained edge devices. This survey provides a comprehensive overview of LLM-SLM collaboration, detailing various interaction mechanisms (pipeline, routing, auxiliary, distillation, fusion), key enabling technologies, and diverse application scenarios driven by on-device needs like low latency, privacy, personalization, and offline operation. While highlighting the significant potential for creating more efficient, adaptable, and accessible AI, we also discuss persistent challenges including system overhead, inter-model consistency, robust task allocation, evaluation complexity, and security/privacy concerns. Future directions point towards more intelligent adaptive frameworks, deeper model fusion, and expansion into multimodal and embodied AI, positioning LLM-SLM collaboration as a key driver for the next generation of practical and ubiquitous artificial intelligence.
Citations
- LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics
- Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
- Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
- Toward Super Agent System with Hybrid AI Routers
- Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration
- KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations
- Self-Resource Allocation in Multi-Agent LLM Systems
- Hawkeye:Efficient Reasoning with Model Collaboration
- MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
- ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
- Life-Cycle Routing Vulnerabilities of LLM Router
- A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval
- Knowledge-Decoupled Synergetic Learning: An MLLM based Collaborative Approach to Few-shot Multimodal Dialogue Intention Recognition
- Collaborative Stance Detection via Small-Large Language Model Consistency Verification
- Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
- Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
- Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
- Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions
- Improving Consistency in Large Language Models through Chain of Guidance
- Dynamic Low-Rank Sparse Adaptation for Large Language Models
- Scaling Multimodal Search and Recommendation with Small Language Models via Upside-Down Reinforcement Learning
- Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
- MixLLM: Dynamic Routing in Mixed Large Language Models
- Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
- Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents
- CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
- Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems
- Multi-Agent Geospatial Copilots for Remote Sensing Workflows
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- A Survey on Large Language Models for Communication, Network, and Service Management: Application Insights, Challenges, and Future Directions
- Smoothie: Label Free Language Model Routing
- A Contemporary Overview: Trends and Applications of Large Language Models on Mobile Devices
- Hymba: A Hybrid-head Architecture for Small Language Models
- FedCoLLM: A Parameter-Efficient Federated Co-tuning Framework for Large and Small Language Models
- CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
- Improving In-Context Learning with Small Language Model Ensembles
- A Review on Edge Large Language Models: Design, Execution, and Applications
- Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment
- Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
- ProFuser: Progressive Fusion of Large Language Models
- Large Language Models (LLMs) for Semantic Communication in Edge-based IoT Networks
- Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
- LLM for Mobile: An Initial Roadmap
- Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
- Llumnix: Dynamic Scheduling for Large Language Model Serving
- Chain of Agents: Large Language Models Collaborating on Long-Context Tasks
- SLMRec: Distilling Large Language Models into Small for Sequential Recommendation
- Thoughtful Things: Building Human-Centric Smart Devices with Small Language Models
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
- Large Language Models Meet User Interfaces: The Case of Provisioning Feedback
- Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
- Auxiliary task demands mask the capabilities of smaller language models
- CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
- A Survey on Knowledge Distillation of Large Language Models
- Personalized Large Language Models
- ML-Enabled Systems Model Deployment and Monitoring: Status Quo and Problems
- Routoo: Learning to Route to Large Language Models Effectively
- Knowledge Fusion of Large Language Models
- Mutual Enhancement of Large and Small Language Models with Cross-Silo Knowledge Transfer
- A Comprehensive Overview of Large Language Models
- Language Models are Few-Shot Learners
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)
- ViLBias: Detecting and Reasoning about Bias in Multimodal Content
Cited by
Related