Tom Everitt
- Scalable agent alignment via reward modeling: a research direction
2018/11/19 by Jan Leike, Leike, Jan, David Krueger +9 · 67 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- AI Safety Gridworlds
2017/11/27 by Jan Leike, Leike, Jan, Miljan Martic +13 · 1 voice · 14 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #cs.AI #cs.LG
- Shaking the foundations: delusions in sequence models for interaction and control
2021/10/20 by Pedro A. Ortega, Ortega, Pedro A., Markus Kunesch +37 · 2 voices · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.LG
- Death and Suicide in Universal Artificial Intelligence
2016/06/02 by Jarryd Martin, Tom Everitt, Martin, Jarryd +4 · 1 voice · 2 citations
Computer Science · #Computability, Logic, AI Algorithms #Evolutionary Algorithms and Applications #Cellular Automata and Applications
- Reward Tampering Problems and Solutions in Reinforcement Learning: A\n Causal Influence Diagram Perspective
2019/08/13 by Tom Everitt, Everitt, Tom, Marcus Hütter +5 · 17 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
- A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
2024/04/23 by Seliem El-Sayed, Canfer Akbulut, El-Sayed, Seliem +38 · 2 voices · 9 citations
Psychology · Social Sciences · #Mental Health Research Topics #Ethics and Social Impacts of AI
- General agents contain world models
2025/06/02 by Jonathan Richens, Richens, Jonathan, David Abel +5 · 6 voices · 13 citations
#cs.AI #cs.LG #cs.RO #stat.ML
- Robust agents learn causal world models
2024/02/16 by Jonathan G. Richens, Tom Everitt, Richens, Jonathan +1 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Avoiding Wireheading with Value Reinforcement Learning
2016/05/10 by Tom Everitt, Everitt, Tom, Marcus Hütter +1 · 2 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Game Theory and Applications #Reinforcement Learning in Robotics
- Evaluating the Goal-Directedness of Large Language Models
2025/04/16 by Tom Everitt, Cristina Garbacea, Everitt, Tom +12 · 1 voice · 6 citations
Computer Science · #Topic Modeling #Text Readability and Simplification #Natural Language Processing Techniques
- Human Control: Definitions and Algorithms
2023/05/31 by Ryan M. Carey, Carey, Ryan, Tom Everitt +1 · 2 citations
Psychology · #Human-Automation Interaction and Safety
- Self-Modification of Policy and Utility Function in Rational Agents
2016/05/10 by Tom Everitt, Daniel Filan, Everitt, Tom +5 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Reinforcement Learning in Robotics
- Path-Specific Objectives for Safer Agent Incentives
2022/04/21 by Sebastian Farquhar, Ryan M. Carey, Farquhar, Sebastian +3 · 1 citation
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (stat.ML)
- Measuring Goal-Directedness
2024/12/06 by Matt MacDermott, MacDermott, Matt, James Fox +5 · 2 citations
Decision Sciences · #Evaluation and Performance Assessment