2022/05/27 by Ankita Tondwalkar, Tondwalkar, Ankita, Andres Kwasinski +1
Computer Science · Engineering · #Cognitive Radio Networks and Spectrum Sensing #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #Full-Duplex Wireless Communications #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2205.13944
openalex publication_date 2022/05/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper presents a novel deep reinforcement learning-based resource allocation technique for the multi-agent environment presented by a cognitive radio network where the interactions of the agents during learning may lead to a non-stationary environment. The resource allocation technique presented in this work is distributed, not requiring coordination with other agents. It is shown by considering aspects specific to deep reinforcement learning that the presented algorithm converges in an arbitrarily long time to equilibrium policies in a non-stationary multi-agent environment that results from the uncoordinated dynamic interaction between radios through the shared wireless environment. Simulation results show that the presented technique achieves a faster learning performance compared to an equivalent table-based Q-learning algorithm and is able to find the optimal policy in 99% of cases for a sufficiently long learning time. In addition, simulations show that our DQL approach requires less than half the number of learning steps to achieve the same performance as an equivalent table-based implementation. Moreover, it is shown that the use of a standard single-agent deep reinforcement learning approach may not achieve convergence when used in an uncoordinated interacting multi-radio scenario