vix.ing · top · new · best · stats

Deep Actor-Critic Learning for Distributed Power Control in Wireless Mobile Networks

2020/09/14 by Yasar Sinan Nasir, Dongning Guo, Nasir, Yasar Sinan +1
Computer Science · Engineering · Mathematics · #Adaptive Dynamic Programming Control #Cognitive Radio Networks and Spectrum Sensing #Distributed Control Multi-Agent Systems #FOS: Computer and information sciences #FOS: Electrical engineering #Information Theory (cs.IT) #Machine Learning (stat.ML) #Signal Processing (eess.SP) #cs.IT #eess.SP #electronic engineering #information engineering #math.IT #stat.ML

paper · pdf · doi:10.48550/arxiv.2009.06681

5 pages, 4 figures, to appear in the 54th Annual IEEE Asilomar Conference on Signals, Systems, and Computers, Nov 2020. This is an invited paper to the session Reinforcement Learning and Bandits for Communication Systems. To reproduce the results please see https://github.com/sinannasir/Power-Control-asilomar

arxiv created 2020/09/14 · openalex publication_date 2020/09/14 · arxiv updated 2020/09/16 · openalex created_date 2020/09/21 · openalex updated_date 2026/07/28

Abstract

Deep reinforcement learning offers a model-free alternative to supervised deep learning and classical optimization for solving the transmit power control problem in wireless networks. The multi-agent deep reinforcement learning approach considers each transmitter as an individual learning agent that determines its transmit power level by observing the local wireless environment. Following a certain policy, these agents learn to collaboratively maximize a global objective, e.g., a sum-rate utility function. This multi-agent scheme is easily scalable and practically applicable to large-scale cellular networks. In this work, we present a distributively executed continuous power control algorithm with the help of deep actor-critic learning, and more specifically, by adapting deep deterministic policy gradient. Furthermore, we integrate the proposed power control algorithm to a time-slotted system where devices are mobile and channel conditions change rapidly. We demonstrate the functionality of the proposed algorithm using simulation results.

Citations

Related