2022/05/31 by Alexander Tyurin, Peter Richtárik, Tyurin, Alexander +1
Computer Science · #Stochastic Gradient Optimization Techniques #Privacy-Preserving Technologies in Data #Cooperative Communication and Network Coding
paper · pdf · doi:10.48550/arxiv.2205.15580
We present a new method that includes three key components of distributed optimization and federated learning: variance reduction of stochastic gradients, partial participation, and compressed communication. We prove that the new method has optimal oracle complexity and state-of-the-art communication complexity in the partial participation setting. Regardless of the communication compression feature, our method successfully combines variance reduction and partial participation: we get the optimal oracle complexity, never need the participation of all nodes, and do not require the bounded gradients (dissimilarity) assumption.