vix.ing · top · new · best · stats

On Learning to Think: Algorithmic Information Theory for Novel\n Combinations of Reinforcement Learning Controllers and Recurrent Neural World\n Models

2015/11/30 by Juergen Schmidhuber, Schmidhuber, Juergen · 5 voices · 26 citations
Computer Science · Engineering · #Advanced Memory and Neural Computing #Adversarial Robustness in Machine Learning #Artificial intelligence #Artificial neural network #Computer science #Exploit #Machine learning #Neural Networks and Applications #Recurrent neural network #Reinforcement Learning in Robotics #Reinforcement learning #Scratch #cs.AI #cs.LG #cs.NE

paper · pdf · doi:10.48550/arxiv.1511.09249

published in arXiv (Cornell University) (Cornell University) · 36 pages, 1 figure. arXiv admin note: substantial text overlap with arXiv:1404.7828

arxiv created 2015/11/30 · openalex publication_date 2015/11/30 · arxiv updated 2015/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper addresses the general problem of reinforcement learning (RL) in\npartially observable environments. In 2013, our large RL recurrent neural\nnetworks (RNNs) learned from scratch to drive simulated cars from\nhigh-dimensional video input. However, real brains are more powerful in many\nways. In particular, they learn a predictive model of their initially unknown\nenvironment, and somehow use it for abstract (e.g., hierarchical) planning and\nreasoning. Guided by algorithmic information theory, we describe RNN-based AIs\n(RNNAIs) designed to do the same. Such an RNNAI can be trained on never-ending\nsequences of tasks, some of them provided by the user, others invented by the\nRNNAI itself in a curious, playful fashion, to improve its RNN-based world\nmodel. Unlike our previous model-building RNN-based RL machines dating back to\n1990, the RNNAI learns to actively query its model for abstract reasoning and\nplanning and decision making, essentially "learning to think." The basic ideas\nof this report can be applied to many other cases where one RNN-like system\nexploits the algorithmic information content of another. They are taken from a\ngrant proposal submitted in Fall 2014, and also explain concepts such as\n"mirror neurons." Experimental results will be described in separate papers.\n

Citations

Cited by

Discussions

Related