Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
2025/05/26 by Liu, Junming, Gao, Yanting, Meng, Siyuan +6
#Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)
paper · doi:10.48550/arxiv.2505.19699
Abstract
Federated Learning (FL) is a decentralized machine learning paradigm that enables clients to collaboratively train models while preserving data privacy. However, the coexistence of model and data heterogeneity gives rise to inconsistent representations and divergent optimization dynamics across clients, ultimately hindering robust global performance. To transcend these challenges, we propose Mosaic, a novel data-free knowledge distillation framework tailored for heterogeneous distributed environments. Mosaic first trains local generative models to approximate each client's personalized distribution, enabling synthetic data generation that safeguards privacy through strict separation from real data. Subsequently, Mosaic forms a Mixture-of-Experts (MoE) from client models based on their specialized knowledge, and distills it into a global model using the generated data. To further enhance the MoE architecture, Mosaic integrates expert predictions via a lightweight meta model trained on a few representative prototypes. Extensive experiments on standard image classification benchmarks demonstrate that Mosaic consistently outperforms state-of-the-art approaches under both model and data heterogeneity. The source code has been published at https://github.com/Wings-Of-Disaster/Mosaic.
Citations
- Boosting Generalization Performance in Model-Heterogeneous Federated Learning Using Variational Transposed Convolution
- A Scalable Unsupervised Framework for multi-aspect labeling of Multilingual and Multi-Domain Review Data
- ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
- Personalized Federated Learning via Feature Distribution Adaptation
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion
- Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
- DFDG: Data-Free Dual-Generator Adversarial Distillation for One-Shot Federated Learning
- Privacy-Preserving Federated Learning with Consistency via Knowledge Distillation Using Conditional Generator
- Threats and Defenses in Federated Learning Life Cycle: A Comprehensive Survey and Challenges
- An Aggregation-Free Federated Learning for Tackling Data Heterogeneity
- FedTAD: Topology-aware Data-free Knowledge Distillation for Subgraph Federated Learning
- Federated Distillation: A Survey
- De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
- FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Data-free Knowledge Distillation for Fine-grained Visual Categorization
- NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated Learning
- Federated Learning for Connected and Automated Vehicles: A Survey of Existing Approaches and Challenges
- Towards Personalized Federated Learning via Heterogeneous Model Reassembly
- ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting
- Heterogeneous Federated Learning: State-of-the-art and Research Challenges
- Towards Open Federated Learning Platforms: Survey and Vision from Technical and Legal Perspectives
- FedMultimodal: A Benchmark For Multimodal Federated Learning
- Understanding How Consistency Works in Federated Learning via Stage-wise Relaxed Initialization
- Tackling Data Heterogeneity in Federated Learning with Class Prototypes
- FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction
- Diffusion Models for Medical Image Analysis: A Comprehensive Survey
- Client Selection in Federated Learning: Principles, Challenges, and Opportunities
- A Survey on Heterogeneous Federated Learning
- Rethinking Data Heterogeneity in Federated Learning: Introducing a New Notion and Standard Benchmarks
- Virtual Homogeneity Learning: Defending against Data Heterogeneity in Federated Learning
- FEDIC: Federated Learning on Non-IID and Long-Tailed Data via Calibrated Distillation
- Federated Learning on Heterogeneous and Long-Tailed Data via Classifier Re-Training with Federated Features
- FedDC: Federated Learning with Non-IID Data via Local Drift Decoupling and Correction
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning
- Architecture Agnostic Federated Learning for Neural Networks
- Federated Learning of Generative Image Priors for MRI Reconstruction
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay
- FedDTG:Federated Data-Free Knowledge Distillation via Three-Player Generative Adversarial Networks
- DENSE: Data-Free One-Shot Federated Learning
- Local Learning Matters: Rethinking Data Heterogeneity in Federated Learning
- Federated Learning for Smart Healthcare: A Survey
- Deceive D: Adaptive Pseudo Augmentation for GAN Training with Limited Data
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- Asynchronous Federated Learning on Heterogeneous Devices: A Survey
- Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning
- FedProto: Federated Prototype Learning across Heterogeneous Clients
- FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks
- Model-Contrastive Federated Learning
- Pros and Cons of GAN Evaluation Measures: New Developments
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered Dropout
- Learning Transferable Visual Models From Natural Language Supervision
- FedBN: Federated Learning on Non-IID Features via Local Batch Normalization
- Differentially Private Secure Multi-Party Computation for Federated\n Learning in Financial Applications
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients
- Federated Mutual Learning
- Denoising Diffusion Probabilistic Models
- FedGAN: Federated Generative Adversarial Networks for Distributed Data
- Knowledge Distillation: A Survey
- Training Generative Adversarial Networks with Limited Data
- LDP-Fed: Federated Learning with Local Differential Privacy
- Adaptive Federated Optimization
- Federated Learning with Matched Averaging
- A Simple Framework for Contrastive Learning of Visual Representations
- DeGAN : Data-Enriching GAN for Retrieving Representative Samples from a\n Trained Classifier
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- Advances and Open Problems in Federated Learning
- Federated Learning with Differential Privacy: Algorithms and Performance Analysis
- Federated Learning With Differential Privacy: Algorithms and Performance Analysis
- FedMD: Heterogenous Federated Learning via Model Distillation
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Federated Learning: Challenges, Methods, and Future Directions
- Bayesian Nonparametric Federated Learning of Neural Networks
- One-Shot Federated Learning
- Federated Optimization in Heterogeneous Networks
- Expanding the Reach of Federated Learning by Reducing Client Resource Requirements
- CrisisMMD: Multimodal Twitter Datasets from Natural Disasters
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Attention Is All You Need
- Unrolled Generative Adversarial Networks
- Membership Inference Attacks against Machine Learning Models
- Improved Techniques for Training GANs
- Deep Residual Learning for Image Recognition
- You Only Look Once: Unified, Real-Time Object Detection
- Distilling the Knowledge in a Neural Network
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Fully Convolutional Networks for Semantic Segmentation
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ImageNet classification with deep convolutional neural networks
- Stacked generalization
- Adaptive Mixtures of Local Experts
- Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
- Gradient-based learning applied to document recognition
Related