2021/02/08 by Rahul Meshram, Meshram, Rahul, Kesav Kaza +1 · 1 citation
Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Energy Load and Power Forecasting #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Smart Grid Energy Management #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2102.04321
openalex publication_date 2021/02/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We model online recommendation systems using the hidden Markov multi-state\nrestless multi-armed bandit problem. To solve this we present Monte Carlo\nrollout policy. We illustrate numerically that Monte Carlo rollout policy\nperforms better than myopic policy for arbitrary transition dynamics with no\nspecific structure. But, when some structure is imposed on the transition\ndynamics, myopic policy performs better than Monte Carlo rollout policy.\n