2021/07/25 by Howard H. Yang, Zihan Chen, Yang, Howard H. +5 · 4 citations
Computer Science · Engineering · #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2107.11733
openalex publication_date 2021/07/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study a distributed machine learning problem carried out by an edge server\nand multiple agents in a wireless network. The objective is to minimize a\nglobal function that is a sum of the agents' local loss functions. And the\noptimization is conducted by analog over-the-air model training. Specifically,\neach agent modulates its local gradient onto a set of waveforms and transmits\nto the edge server simultaneously. From the received analog signal the edge\nserver extracts a noisy aggregated gradient which is distorted by the channel\nfading and interference, and uses it to update the global model and feedbacks\nto all the agents for another round of local computing. Since the\nelectromagnetic interference generally exhibits a heavy-tailed intrinsic, we\nuse the \α-stable distribution to model its statistic. In consequence,\nthe global gradient has an infinite variance that hinders the use of\nconventional techniques for convergence analysis that rely on second-order\nmoments' existence. To circumvent this challenge, we take a new route to\nestablish the analysis of convergence rate, as well as generalization error, of\nthe algorithm. Our analyses reveal a two-sided effect of the interference on\nthe overall training procedure. On the negative side, heavy tail noise slows\ndown the convergence rate of the model training: the heavier the tail in the\ndistribution of interference, the slower the algorithm converges. On the\npositive side, heavy tail noise has the potential to increase the\ngeneralization power of the trained model: the heavier the tail, the better the\nmodel generalizes. This perhaps counterintuitive conclusion implies that the\nprevailing thinking on interference -- that it is only detrimental to the edge\nlearning system -- is outdated and we shall seek new techniques that exploit,\nrather than simply mitigate, the interference for better machine learning in\nwireless networks.\n