vix.ing · top · new · best · stats · spec

Offline Reinforcement Learning from Human Feedback in Real-World\n Sequence-to-Sequence Tasks

2020/11/04 by Julia Kreutzer, Kreutzer, Julia, Stefan Riezler +3 · 1 citation
Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Flexible and Reconfigurable Manufacturing Systems #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Scheduling and Optimization Algorithms

paper · pdf · doi:10.48550/arxiv.2011.02511

openalex publication_date 2020/11/04 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Large volumes of interaction logs can be collected from NLP systems that are\ndeployed in the real world. How can this wealth of information be leveraged?\nUsing such interaction logs in an offline reinforcement learning (RL) setting\nis a promising approach. However, due to the nature of NLP tasks and the\nconstraints of production systems, a series of challenges arise. We present a\nconcise overview of these challenges and discuss possible solutions.\n

Cited by

Related