2017/09/01 by Jan Oksanen, Visa Koivunen, Oksanen, Jan +1
Computer Science · Decision Sciences · Economics, Econometrics and Finance · Engineering · #Advanced Bandit Algorithms Research #Cognitive Radio Networks and Spectrum Sensing #FOS: Computer and information sciences #FOS: Electrical engineering #Financial Markets and Investment Strategies #Information Theory (cs.IT) #Signal Processing (eess.SP) #Smart Grid Energy Management #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1709.00237
openalex publication_date 2017/09/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper a spectrum sensing policy employing recency-based exploration\nis proposed for cognitive radio networks. We formulate the problem of finding a\nspectrum sensing policy for multi-band dynamic spectrum access as a stochastic\nrestless multi-armed bandit problem with stationary unknown reward\ndistributions. In cognitive radio networks the multi-armed bandit problem\narises when deciding where in the radio spectrum to look for idle frequencies\nthat could be efficiently exploited for data transmission. We consider two\nmodels for the dynamics of the frequency bands: 1) the independent model where\nthe state of the band evolves randomly independently from the past and 2) the\nGilbert-Elliot model, where the states evolve according to a 2-state Markov\nchain. It is shown that in these conditions the proposed sensing policy attains\nasymptotically logarithmic weak regret. The policy proposed in this paper is an\nindex policy, in which the index of a frequency band is comprised of a sample\nmean term and a recency-based exploration bonus term. The sample mean promotes\nspectrum exploitation whereas the exploration bonus encourages for further\nexploration for idle bands providing high data rates. The proposed recency\nbased approach readily allows constructing the exploration bonus such that it\nwill grow the time interval between consecutive sensing time instants of a\nsuboptimal band exponentially, which then leads to logarithmically increasing\nweak regret. Simulation results confirming logarithmic weak regret are\npresented and it is found that the proposed policy provides often improved\nperformance at low complexity over other state-of-the-art policies in the\nliterature.\n