2011/02/23 by Xu Chen, Chen, Xu, Jianwei Huang +3
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Advanced MIMO Systems Optimization #Cognitive Radio Networks and Spectrum Sensing #Distributed #FOS: Computer and information sciences #Networking and Internet Architecture (cs.NI) #Parallel #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.1102.4728
openalex publication_date 2011/02/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a dynamic spectrum access scheme where secondary users recommend "good" channels to each other and access accordingly. We formulate the problem as an average reward based Markov decision process. We show the existence of the optimal stationary spectrum access policy, and explore its structure properties in two asymptotic cases. Since the action space of the Markov decision process is continuous, it is difficult to find the optimal policy by simply discretizing the action space and use the policy iteration, value iteration, or Q-learning methods. Instead, we propose a new algorithm based on the Model Reference Adaptive Search method, and prove its convergence to the optimal policy. Numerical results show that the proposed algorithms achieve up to 18% and 100% performance improvement than the static channel recommendation scheme in homogeneous and heterogeneous channel environments, respectively, and is more robust to channel dynamics.