vix.ing · top · new · best · stats · spec

Accelerated Structure-Aware Reinforcement Learning for Delay-Sensitive\n Energy Harvesting Wireless Sensors

2018/07/22 by Nikhilesh Sharma, Sharma, Nikhilesh, Nicholas Mastronarde +3 · 1 citation
Computer Science · Engineering · #Advanced MIMO Systems Optimization #Age of Information Optimization #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Networking and Internet Architecture (cs.NI) #Signal Processing (eess.SP) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1807.08315

openalex publication_date 2018/07/22 · openalex created_date 2022/08/04 · openalex updated_date 2026/07/28

Abstract

We investigate an energy-harvesting wireless sensor transmitting\nlatency-sensitive data over a fading channel. The sensor injects captured data\npackets into its transmission queue and relies on ambient energy harvested from\nthe environment to transmit them. We aim to find the optimal scheduling policy\nthat decides whether or not to transmit the queue's head-of-line packet at each\ntransmission opportunity such that the expected packet queuing delay is\nminimized given the available harvested energy. No prior knowledge of the\nstochastic processes that govern the channel, captured data, or harvested\nenergy dynamics are assumed, thereby necessitating the use of online learning\nto optimize the scheduling policy. We formulate this scheduling problem as a\nMarkov decision process (MDP) and analyze the structural properties of its\noptimal value function. In particular, we show that it is non-decreasing and\nhas increasing differences in the queue backlog and that it is non-increasing\nand has increasing differences in the battery state. We exploit this structure\nto formulate a novel accelerated reinforcement learning (RL) algorithm to solve\nthe scheduling problem online at a much faster learning rate, while limiting\nthe induced computational complexity. Our experiments demonstrate that the\nproposed algorithm closely approximates the performance of an optimal offline\nsolution that requires a priori knowledge of the channel, captured data, and\nharvested energy dynamics. Simultaneously, by leveraging the value function's\nstructure, our approach achieves competitive performance relative to a\nstate-of-the-art RL algorithm, at potentially orders of magnitude lower\ncomplexity. Finally, considerable performance gains are demonstrated over the\nwell-known and widely used Q-learning algorithm.\n

Cited by

Related