Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
2025/04/24 by Yang Liu, Liu, Yang, Zhang, Kejia +22 · 8 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model-Driven Software Engineering Techniques #Scientific Computing and Data Management
paper · pdf · doi:10.48550/arxiv.2504.17421
openalex publication_date 2025/04/24 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28
Abstract
Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain-specific tasks; small models (SMs), in contrast, are more efficient and tailored to specific domains yet lack general-purpose coverage. Taking a collaborative approach, where large and small models work synergistically, can accelerate the adaptation of LLMs to private domains and unlock new potential in AI. This survey presents a comprehensive overview of recent advances and challenges in harnessing the collaborative power of large and small models for private-domain adaptation. It specifically focuses on the unique constraints of cross-boundary environments, where models belong to distinct parties, and examines the resulting tensions among data privacy, model security, integrity, and resource limitations. By analyzing the information flow between distinct model and data stakeholders, we propose a unified taxonomy that classifies research into three primary directions: downward knowledge transfer (LM to SM), upward knowledge transfer (SM to LM), and inference-time collaboration across parties. Drawing on this taxonomy, we analyze the core challenges inherent to cross-boundary information exchange, including data-privacy, model-security, and integrity threats as well as efficiency constraints, and synthesize these into a multi-objective optimization problem that governs practical deployment. Finally, we review key open challenges inherent to such hybrid approaches and outline promising directions for future research. By offering a principled, boundary-centric view of this rapidly evolving landscape, this survey aims to serve as a structured resource for researchers and practitioners advancing privacy-aware, resource-efficient AI deployment.
Citations
- On the Privacy Risk of In-context Learning
- Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
- FuseChat: Knowledge Fusion of Chat Models
- Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
- Casper: Prompt Sanitization for Protecting User Privacy in Web-Based Large Language Models
- LLAVADI: What Matters For Multimodal Large Language Models Distillation
- SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
- DDK: Distilling Domain Knowledge for Efficient Large Language Models
- ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
- Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
- FedBiOT: LLM Local Fine-tuning in Federated Learning without Full Model
- FuseGen: PLM Fusion for Data-generation based Zero-shot Learning
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
- FedMKT: Federated Mutual Knowledge Transfer for Large and Small Language Models
- Efficient Adversarial Training in LLMs with Continuous Attacks
- Federated Adaptation for Foundation Model-based Recommendations
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Offset Unlearning for Large Language Models
- FedPFT: Federated Proxy Fine-Tuning of Foundation Models
- Privacy Preserving Prompt Engineering: A Survey
- BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models
- An Upload-Efficient Scheme for Transferring Knowledge From a Server-Side Pre-trained Generator to Clients in Heterogeneous Federated Learning
- CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
- Differentially Private Synthetic Data via Foundation Model APIs 2: Text
- The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
- Privacy-Preserving Instructions for Aligning Large Language Models
- A Survey on Knowledge Distillation of Large Language Models
- Privacy-Preserving Language Model Inference with Instance Obfuscation
- History, Development, and Principles of Large Language Models-An Introductory Survey
- MEA-Defender: A Robust Watermark against Model Extraction Attack
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- Knowledge Fusion of Large Language Models
- Tuning Language Models by Proxy
- A Survey of Resource-efficient LLM and Multimodal Foundation Models
- Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
- FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning
- A Split-and-Privatize Framework for Large Language Model Fine-Tuning
- Cascade Speculative Drafting for Even Faster LLM Inference
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Mutual Enhancement of Large and Small Language Models with Cross-Silo Knowledge Transfer
- Scalable Extraction of Training Data from (Production) Language Models
- DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer
- Tunable Soft Prompts are Messengers in Federated Learning
- Orchestration of Emulator Assisted Mobile Edge Tuning for AI Foundation Models: A Multi-Agent Deep Reinforcement Learning Approach
- Retrieval-based Knowledge Transfer: An Effective Approach for Extreme Large Language Model Compression
- CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language Model
- SpecTr: Fast Speculative Decoding via Optimal Transport
- Let's Synthesize Step by Step: Iterative Dataset Synthesis with Large Language Models by Extrapolating Errors from Small Models
- An Emulator for Fine-Tuning Large Language Models using Small Language Models
- Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric Perspective
- FATE-LLM: A Industrial Grade Federated Learning Framework for Large Language Models
- Online Speculative Decoding
- Efficient Federated Prompt Tuning for Black-box Large Pre-trained Models
- Model Leeching: An Extraction Attack Targeting LLMs
- Hide and Seek (HaS): A Lightweight Framework for Prompt Privacy Protection
- FederatedScope-LLM: A Comprehensive Package for Fine-tuning Large Language Models in Federated Learning
- Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty
- Reverse Knowledge Distillation: Training a Large Model using a Small One for Retinal Image Matching on Limited Data
- Composing Parameter-Efficient Modules with Arithmetic Operations
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- A Simple and Effective Pruning Approach for Large Language Models
- Bridging the Gap between Decision and Logits in Decision-based Knowledge Distillation for Pre-trained Language Models
- Protecting User Privacy in Remote Conversational Systems: A Privacy-Preserving framework based on text sanitization
- Harnessing large-language models to generate private synthetic text
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Training Data Extraction From Pre-trained Language Models: A Survey
- Differentially Private Synthetic Data via Foundation Model APIs 1: Images
- Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models
- CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models
- Can Public Large Language Models Help Private Cross-device Federated Learning?
- Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
- Breaching FedMD: Image Recovery via Paired-Logits Inversion Attack
- BloombergGPT: A Large Language Model for Finance
- BlackVIP: Black-Box Visual Prompting for Robust Transfer Learning
- FedML-HE: An Efficient Homomorphic-Encryption-Based Privacy-Preserving Federated Learning System
- Efficient and Secure Federated Learning for Financial Applications
- GPT-4 Technical Report
- A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
- Multimodal Federated Learning via Contrastive Representation Ensemble
- Speculative Decoding with Big Little Decoder
- Offsite-Tuning: Transfer Learning without Full Model
- Use of Federated Learning and Blockchain towards Securing Financial Services
- REPLUG: Retrieval-Augmented Black-Box Language Models
- When Federated Learning Meets Pre-trained Language Models' Parameter-Efficient Tuning Methods
- Dataless Knowledge Fusion by Merging Weights of Language Models
- Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
- Fast Inference from Transformers via Speculative Decoding
- Vertical Federated Learning: Concepts, Advances, and Challenges
- Contrastive Decoding: Open-ended Text Generation as Optimization
- Will we run out of data? Limits of LLM scaling based on human-generated data
- ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback
- Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering
- Knowledge Unlearning for Mitigating Privacy Risks in Language Models
- Less is More: Task-aware Layer-wise Distillation for Language Model Compression
- Federated Learning from Pre-Trained Models: A Contrastive Learning Approach
- YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications
- Trading Off Privacy, Utility and Efficiency in Federated Learning
- FedPrompt: Communication-Efficient and Privacy Preserving Prompt Tuning in Federated Learning
- Federated Learning via Decentralized Dataset Distillation in Resource-Constrained Edge Environments
- YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors
- Preserving Privacy in Federated Learning with Ensemble Cross-Domain Knowledge Distillation
- Federated Learning with GAN-based Data Synthesis for Non-IID Clients
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
- Privacy for Free: How does Dataset Condensation Help Privacy?
- THE-X: Privacy-Preserving Transformer Inference with Homomorphic Encryption
- Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
- IDEAL: Query-Efficient Data-Free Learning from Black-box Models
- Towards Data-Free Model Stealing in a Hard Label Setting
- ZeroGen: Efficient Zero-shot Learning via Dataset Generation
- Generating Training Data with Language Models: Towards Zero-Shot Language Understanding
- The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
- The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Parameterized Knowledge Transfer for Personalized Federated Learning
- FedGEMS: Federated Learning of Larger Server Models via Selective Knowledge Fusion
- Learning to Teach with Student Feedback
- CAPE: Context-Aware Private Embeddings for Private Language Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- BERT Learns to Teach: Knowledge Distillation with Meta Learning
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box Model
- FedProto: Federated Prototype Learning across Heterogeneous Clients
- See through Gradients: Image Batch Recovery via GradInversion
- Natural Language Understanding with Privacy-Preserving BERT
- Extracting Training Data from Large Language Models
- Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness
- Distilled One-Shot Federated Learning
- Breaking the Communication-Privacy-Accuracy Trilemma
- Towards Differentially Private Text Representations
- Mix2FLD: Downlink Federated Learning After Uplink Federated Distillation With Two-Way Mixup
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Dataset Condensation with Gradient Matching
- Knowledge Distillation: A Survey
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Information Leakage in Embedding Models
- Meta Pseudo Labels
- Advances and Open Problems in Federated Learning
- FedMD: Heterogenous Federated Learning via Model Distillation
- Deep Leakage from Gradients
- Towards Federated Learning at Scale: System Design
- BioBERT: a pre-trained biomedical language representation model for biomedical text mining
- Split learning for health: Distributed deep learning without sharing raw patient data
- Privacy-preserving Neural Representations of Text
- Towards Robust and Privacy-preserving Text Representations
- YOLOv3: An Incremental Improvement
- mixup: Beyond Empirical Risk Minimization
- YOLO9000: Better, Faster, Stronger
- Deep Learning with Differential Privacy
- Communication-Efficient Learning of Deep Networks from Decentralized Data
- Deep Residual Learning for Image Recognition
- You Only Look Once: Unified, Real-Time Object Detection
- Distilling the Knowledge in a Neural Network
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Calibrating Noise to Sensitivity in Private Data Analysis
- Long Short-Term Memory
- LatticeGen: A Cooperative Framework which Hides Generated Text in a Lattice for Privacy-Aware Generation on Cloud
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
Cited by
Related