vix.ing · top · new · best · stats

Reinforcement Learning for Bandit Neural Machine Translation with\n Simulated Human Feedback

2017/07/24 by Khanh Nguyen, Nguyen, Khanh, Hal Daumé +5 · 7 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Artificial intelligence #Artificial neural network #Computation and Language (cs.CL) #Computer science #Encoder #Engineering #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine learning #Machine translation #Reinforcement #Reinforcement learning #Topic Modeling #Translation (biology) #cs.AI #cs.CL #cs.HC #cs.LG

paper · pdf · doi:10.48550/arxiv.1707.07402

published in arXiv (Cornell University) (Cornell University) · 11 pages, 5 figures, In Proceedings of Empirical Methods in Natural Language Processing (EMNLP) 2017

openalex publication_date 2017/07/24 · arxiv created 2017/11/11 · arxiv updated 2017/11/15 · openalex created_date 2022/10/04 · openalex updated_date 2026/08/06

Abstract

Machine translation is a natural candidate problem for reinforcement learning\nfrom human feedback: users provide quick, dirty ratings on candidate\ntranslations to guide a system to improve. Yet, current neural machine\ntranslation training focuses on expensive human-generated reference\ntranslations. We describe a reinforcement learning algorithm that improves\nneural machine translation systems from simulated human feedback. Our algorithm\ncombines the advantage actor-critic algorithm (Mnih et al., 2016) with the\nattention-based neural encoder-decoder architecture (Luong et al., 2015). This\nalgorithm (a) is well-designed for problems with a large action space and\ndelayed rewards, (b) effectively optimizes traditional corpus-level machine\ntranslation metrics, and (c) is robust to skewed, high-variance, granular\nfeedback modeled after actual human behaviors.\n

Cited by

Related