Adaptive Federated Optimization
2020/02/29 by Sashank Reddi, Sashank J. Reddi, Zachary Charles +14 · 1 voice · 148 citations
Computer Science · Mathematics · #Distributed #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Optimization and Control (math.OC) #Parallel #Privacy-Preserving Technologies in Data #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC) #cs.DC #cs.LG #math.OC #stat.ML
paper · pdf · doi:10.48550/arxiv.2003.00295
Published as a conference paper at ICLR 2021
openalex publication_date 2020/02/29 · arxiv published 2020/02/29 · arxiv created 2021/09/08 · arxiv updated 2021/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Federated learning is a distributed machine learning paradigm in which a large number of clients coordinate with a central server to learn a model without sharing their own training data. Standard federated optimization methods such as Federated Averaging (FedAvg) are often difficult to tune and exhibit unfavorable convergence behavior. In non-federated settings, adaptive optimization methods have had notable success in combating such issues. In this work, we propose federated versions of adaptive optimizers, including Adagrad, Adam, and Yogi, and analyze their convergence in the presence of heterogeneous data for general non-convex settings. Our results highlight the interplay between client heterogeneity and communication efficiency. We also perform extensive experiments on these methods and show that the use of adaptive optimizers can significantly improve the performance of federated learning.
Citations
Cited by
- First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
- Clust-PSI-PFL: A Population Stability Index Approach for Clustered Non-IID Personalized Federated Learning
- FedDPC : Handling Data Heterogeneity and Partial Client Participation in Federated Learning
- FedPOD: the deployable units of training for federated learning
- Learned Digital Codes for Over-the-Air Computation in Federated Edge Learning
- Privacy-Enhancing Infant Cry Classification with Federated Transformers and Denoising Regularization
- Adaptive federated learning for ship detection across diverse satellite imagery sources
- Semantic-Constrained Federated Aggregation: Convergence Theory and Privacy-Utility Bounds for Knowledge-Enhanced Distributed Learning
- FedLAD: A Modular and Adaptive Testbed for Federated Log Anomaly Detection
- An Accelerated Primal Dual Algorithm with Backtracking for Decentralized Constrained Optimization
- Geometric Prior-Guided Federated Prompt Calibration
- MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
- Stragglers Can Contribute More: Uncertainty-Aware Distillation for Asynchronous Federated Learning
- ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
- On the Limits of Momentum in Decentralized and Federated Optimization
- Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
- CycleSL: Server-Client Cyclical Update Driven Scalable Split Learning
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning
- A Unified Convergence Analysis for Semi-Decentralized Learning: Sampled-to-Sampled vs. Sampled-to-All Communication
- SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data
- Learning Performance Optimization for Edge AI System with Time and Energy Constraints
- Nesterov-Accelerated Robust Federated Learning Over Byzantine Adversaries
- Split Learning-Enabled Framework for Secure and Light-weight Internet of Medical Things Systems
- Why Federated Optimization Fails to Achieve Perfect Fitting? A Theoretical Perspective on Client-Side Optima
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
- FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
- Privacy-Aware Federated nnU-Net for ECG Page Digitization
- ADP-VRSGP: Decentralized Learning with Adaptive Differential Privacy via Variance-Reduced Stochastic Gradient Push
- FedGPS: Statistical Rectification Against Data Heterogeneity in Federated Learning
- Efficient Multi-Worker Selection based Distributed Swarm Learning via Analog Aggregation
- Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
- Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- FLAMMABLE: A Multi-Model Federated Learning Framework with Multi-Model Engagement and Adaptive Batch Sizes
- Decoupled DiLoCo for Resilient Distributed Pre-training
- FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
- DPMM-CFL: Clustered Federated Learning via Dirichlet Process Mixture Model Nonparametric Clustering
- DP-HYPE: Distributed Differentially Private Hyperparameter Search
- MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
- Adaptive Federated Learning via Dynamical System Model
- FTTE: Enabling Federated and Resource-Constrained Deep Edge Intelligence
- On Provable Benefits of Muon in Federated Learning
- Distributed Low-Communication Training with Decoupled Momentum Optimization
- FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
- TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
- Differential Privacy in Federated Learning: Mitigating Inference Attacks with Randomized Response
- ParaAegis: Parallel Protection for Flexible Privacy-preserved Federated Learning
- Dissecting Federated-Graph Aggregation under Domain Shift: Importance-Aware Aggregation via Empirical Analysis
- FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
- Taming Volatility: Stable and Private QUIC Classification with Federated Learning
- Anomaly Detection in Electric Vehicle Charging Stations Using Federated Learning
- Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
- On Transferring, Merging, and Splitting Task-Oriented Network Digital Twins
- One-Shot Clustering for Federated Learning Under Clustering-Agnostic Assumption
- FediLoRA: Practical Federated Fine-Tuning of Foundation Models Under Missing-Modality Constraints
- Federated Learning for Large Models in Medical Imaging: A Comprehensive Review
- A Study of Privacy-preserving Language Modeling Approaches
- FedGreed: A Byzantine-Robust Loss-Based Aggregation Method for Federated Learning
- FedEve: On Bridging the Client Drift and Period Drift for Cross-device Federated Learning
- Communication-Efficient Federated Learning with Adaptive Number of Participants
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- Robust Federated Learning under Adversarial Attacks via Loss-Based Client Clustering
- Widening the Network Mitigates the Impact of Data Heterogeneity on FedAvg
- On-Device Multimodal Federated Learning for Efficient Jamming Detection
- Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
- HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning
- Decoupled Contrastive Learning for Federated Learning
- A Parameter-free Decentralized Algorithm for Composite Convex Optimization
- Adaptive Stepsize Selection in Decentralized Convex Optimization
- OptiGradTrust: Byzantine-Robust Federated Learning with Multi-Feature Gradient Analysis and Reinforcement Learning-Based Trust Weighting
- Federated Split Learning with Improved Communication and Storage Efficiency
- Optimizing Federated Learning Configurations for MRI Prostate Segmentation and Cancer Detection: A Simulation Study
- FedSWA: Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight Averaging
- Scaling Decentralized Learning with FLock
- FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios
- FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
- A Multi-Objective Optimization framework for Decentralized Learning with coordination constraints
- Federated Learning for Commercial Image Sources
- FLsim: A Modular and Library-Agnostic Simulation Framework for Federated Learning
- Convergence of Agnostic Federated Averaging
- Model Parallelism With Subnetwork Data Parallelism
- Efficient Federated Learning with Timely Update Dissemination
- Towards fair decentralized benchmarking of healthcare AI algorithms with the Federated Tumor Segmentation (FeTS) challenge
- Kalman Filter Aided Federated Koopman Learning
- Towards Decentralized and Sustainable Foundation Model Training with the Edge
- FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
- FedRef: Communication-Efficient Bayesian Fine-Tuning using a Reference Model
- A Practical and Secure Byzantine Robust Aggregator
- FedCLAM: Client Adaptive Momentum with Foreground Intensity Matching for Federated Medical Image Segmentation
- WallStreetFeds: Client-Specific Tokens as Investment Vehicles in Federated Learning
- Federated Loss Exploration for Improved Convergence on Non-IID Data
- GradualDiff-Fed: A Federated Learning Specialized Framework for Large Language Model
- FedCGD: Collective Gradient Divergence Optimized Scheduling for Wireless Federated Learning
- Centroid Approximation for Byzantine-Tolerant Federated Learning
- Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
- AFBS:Buffer Gradient Selection in Semi-asynchronous Federated Learning
- PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
- Decentralized Optimization with Amplified Privacy via Efficient Communication
- FedMLAC: Mutual Learning Driven Heterogeneous Federated Audio Classification
- FedShield-LLM: A Secure and Scalable Federated Fine-Tuned Large Language Model
- Converge Faster, Talk Less: Hessian-Informed Federated Zeroth-Order Optimization
- FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting
- FedRPCA: Enhancing Federated LoRA Aggregation Using Robust PCA
- PSI-PFL: Population Stability Index for Client Selection in non-IID Personalized Federated Learning
- ByzFL: Research Framework for Robust Federated Learning
- Federated Foundation Model for GI Endoscopy Images
- Towards Unified Modeling in Federated Multi-Task Learning via Subspace Decoupling
- Accelerated Training of Federated Learning via Second-Order Methods
- Incorporating Preconditioning into Accelerated Approaches: Theoretical Guarantees and Practical Improvement
- Personalized Subgraph Federated Learning with Differentiable Auxiliary Projections
- MuLoCo: Muon is a practical inner optimizer for DiLoCo
- FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient Estimation
- Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI
- Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
- Incentivizing Permissionless Distributed Learning of LLMs
- Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms
- Federated Instrumental Variable Analysis via Federated Generalized Method of Moments
- Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
- FedSKC: Federated Learning with Non-IID Data via Structural Knowledge Collaboration
- Adaptive Serverless Learning
- ATR-Bench: A Federated Learning Benchmark for Adaptation, Trust, and Reasoning
- RIFLES: Resource-effIcient Federated LEarning via Scheduling
- FedDuA: Doubly Adaptive Federated Learning
- Approximated Behavioral Metric-based State Projection for Federated Reinforcement Learning
- Ranking-Based At-Risk Student Prediction Using Federated Learning and Differential Features
- Adaptive Latent-Space Constraints in Personalized Federated Learning
- Bant: Byzantine Antidote via Trial Function and Trust Scores
- MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning
- Federated generative event models for tokenized electronic health records
- PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs
- Federated learning, ethics, and the double black box problem in medical AI
- Analysis of Asynchronous Federated Learning: Unraveling the Interactions between Gradient Compression, Delay, and Data Heterogeneity
- Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation
- UnifyFL: Enabling Decentralized Cross-Silo Federated Learning
- Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
- Federated Unbiased Learning to Rank
- EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
- SparsyFed: Sparse Adaptive Federated Training
- Private Federated Learning using Preference-Optimized Synthetic Data
- AdGT: Decentralized Gradient Tracking with Tuning-free Per-Agent Stepsize
- Corrected with the Latest Version: Make Robust Asynchronous Federated Learning Possible
- DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
- FedDiverse: Tackling Data Heterogeneity in Federated Learning with Diversity-Driven Client Selection
- FedRecon: Missing Modality Reconstruction in Heterogeneous Distributed Environments
- Accelerating Differentially Private Federated Learning via Adaptive Extrapolation
- Federated Unlearning Made Practical: Seamless Integration via Negated Pseudo-Gradients
- FedFeat+: A Robust Federated Learning Framework Through Federated Aggregation and Differentially Private Feature-Based Classifier Retraining
Discussions
Related