2016/05/30 by Junhyuk Oh, Valliappa Chockalingam, Oh, Junhyuk +5 · 172 citations
Computer Science · Engineering · #Action (physics) #Active perception #Adaptive Dynamic Programming Control #Artificial Intelligence (cs.AI) #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Generalization #Human–computer interaction #Machine Learning (cs.LG) #Observability #Perception #Programming language #Reinforcement Learning in Robotics #Reinforcement learning #Robot #Robotic Locomotion and Control #Set (abstract data type) #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.1605.09128
published in arXiv (Cornell University), 2790-2799 (Cornell University) · ICML 2016
arxiv created 2016/05/30 · openalex publication_date 2016/05/30 · arxiv updated 2016/05/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we introduce a new set of reinforcement learning (RL) tasks in Minecraft (a flexible 3D world). We then use these tasks to systematically compare and contrast existing deep reinforcement learning (DRL) architectures with our new memory-based DRL architectures. These tasks are designed to emphasize, in a controllable manner, issues that pose challenges for RL methods including partial observability (due to first-person visual observations), delayed rewards, high-dimensional visual observations, and the need to use active perception in a correct manner so as to perform well in the tasks. While these tasks are conceptually simple to describe, by virtue of having all of these challenges simultaneously they are difficult for current DRL architectures. Additionally, we evaluate the generalization performance of the architectures on environments not used during training. The experimental results show that our new architectures generalize to unseen environments better than existing DRL architectures.