2018/08/02 by David Rohde, Rohde, David, Stephen Bonner +8 · 15 citations
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Recommender Systems and Techniques #Smart Grid Energy Management
paper · pdf · doi:10.48550/arxiv.1808.00720
openalex publication_date 2018/08/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recommender Systems are becoming ubiquitous in many settings and take many\nforms, from product recommendation in e-commerce stores, to query suggestions\nin search engines, to friend recommendation in social networks. Current\nresearch directions which are largely based upon supervised learning from\nhistorical data appear to be showing diminishing returns with a lot of\npractitioners report a discrepancy between improvements in offline metrics for\nsupervised learning and the online performance of the newly proposed models.\nOne possible reason is that we are using the wrong paradigm: when looking at\nthe long-term cycle of collecting historical performance data, creating a new\nversion of the recommendation model, A/B testing it and then rolling it out. We\nsee that there a lot of commonalities with the reinforcement learning (RL)\nsetup, where the agent observes the environment and acts upon it in order to\nchange its state towards better states (states with higher rewards). To this\nend we introduce RecoGym, an RL environment for recommendation, which is\ndefined by a model of user traffic patterns on e-commerce and the users\nresponse to recommendations on the publisher websites. We believe that this is\nan important step forward for the field of recommendation systems research,\nthat could open up an avenue of collaboration between the recommender systems\nand reinforcement learning communities and lead to better alignment between\noffline and online performance metrics.\n