vix.ing · top · new · best · stats · spec

Monte Carlo Rollout Policy for Recommendation Systems with Dynamic User\n Behavior

2021/02/08 by Rahul Meshram, Meshram, Rahul, Kesav Kaza +1 · 1 citation
Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Energy Load and Power Forecasting #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Smart Grid Energy Management #Systems and Control (eess.SY) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2102.04321

openalex publication_date 2021/02/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We model online recommendation systems using the hidden Markov multi-state\nrestless multi-armed bandit problem. To solve this we present Monte Carlo\nrollout policy. We illustrate numerically that Monte Carlo rollout policy\nperforms better than myopic policy for arbitrary transition dynamics with no\nspecific structure. But, when some structure is imposed on the transition\ndynamics, myopic policy performs better than Monte Carlo rollout policy.\n

Cited by

Related