2020/08/04 by Thijs Vogels, Vogels, Thijs, Sai Praneeth Karimireddy +3 · 3 citations
Computer Science · Engineering · Mathematics · #Algorithm #Artificial intelligence #Bottleneck #Compression (physics) #Compression ratio #Computer science #Deep learning #Distributed #Domain Adaptation and Few-Shot Learning #Engineering #FOS: Computer and information sciences #FOS: Mathematics #Hyperparameter #Information bottleneck method #Lossy compression #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Machine learning #Mathematics #Optimization and Control (math.OC) #Parallel #Rank (graph theory) #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques #Theoretical computer science #and Cluster Computing (cs.DC) #cs.DC #cs.LG #math.OC #stat.ML
paper · pdf · doi:10.48550/arxiv.2008.01425
published in arXiv (Cornell University) (Cornell University) · To appear in NeurIPS 2020
openalex publication_date 2020/08/04 · arxiv created 2020/10/19 · arxiv updated 2020/10/20 · openalex created_date 2022/07/26 · openalex updated_date 2026/08/08
Lossy gradient compression has become a practical tool to overcome the\ncommunication bottleneck in centrally coordinated distributed training of\nmachine learning models. However, algorithms for decentralized training with\ncompressed communication over arbitrary connected networks have been more\ncomplicated, requiring additional memory and hyperparameters. We introduce a\nsimple algorithm that directly compresses the model differences between\nneighboring workers using low-rank linear compressors applied on model\ndifferences. Inspired by the PowerSGD algorithm for centralized deep learning,\nthis algorithm uses power iteration steps to maximize the information\ntransferred per bit. We prove that our method requires no additional\nhyperparameters, converges faster than prior methods, and is asymptotically\nindependent of both the network and the compression. Out of the box, these\ncompressors perform on par with state-of-the-art tuned compression algorithms\nin a series of deep learning benchmarks.\n