vix.ing · top · new · best · stats · spec

Tom Everitt

  1. Scalable agent alignment via reward modeling: a research direction
    2018/11/19 by Jan Leike, Leike, Jan, David Krueger +9 · 67 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  2. AI Safety Gridworlds
    2017/11/27 by Jan Leike, Leike, Jan, Miljan Martic +13 · 1 voice · 14 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #cs.AI #cs.LG
  3. Shaking the foundations: delusions in sequence models for interaction and control
    2021/10/20 by Pedro A. Ortega, Ortega, Pedro A., Markus Kunesch +37 · 2 voices · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.LG
  4. Death and Suicide in Universal Artificial Intelligence
    2016/06/02 by Jarryd Martin, Tom Everitt, Martin, Jarryd +4 · 1 voice · 2 citations
    Computer Science · #Computability, Logic, AI Algorithms #Evolutionary Algorithms and Applications #Cellular Automata and Applications
  5. Reward Tampering Problems and Solutions in Reinforcement Learning: A\n Causal Influence Diagram Perspective
    2019/08/13 by Tom Everitt, Everitt, Tom, Marcus Hütter +5 · 17 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
    2024/04/23 by Seliem El-Sayed, Canfer Akbulut, El-Sayed, Seliem +38 · 2 voices · 9 citations
    Psychology · Social Sciences · #Mental Health Research Topics #Ethics and Social Impacts of AI
  7. General agents contain world models
    2025/06/02 by Jonathan Richens, Richens, Jonathan, David Abel +5 · 6 voices · 13 citations
    #cs.AI #cs.LG #cs.RO #stat.ML
  8. Robust agents learn causal world models
    2024/02/16 by Jonathan G. Richens, Tom Everitt, Richens, Jonathan +1 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Avoiding Wireheading with Value Reinforcement Learning
    2016/05/10 by Tom Everitt, Everitt, Tom, Marcus Hütter +1 · 2 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Game Theory and Applications #Reinforcement Learning in Robotics
  10. Evaluating the Goal-Directedness of Large Language Models
    2025/04/16 by Tom Everitt, Cristina Garbacea, Everitt, Tom +12 · 1 voice · 6 citations
    Computer Science · #Topic Modeling #Text Readability and Simplification #Natural Language Processing Techniques
  11. Human Control: Definitions and Algorithms
    2023/05/31 by Ryan M. Carey, Carey, Ryan, Tom Everitt +1 · 2 citations
    Psychology · #Human-Automation Interaction and Safety
  12. Self-Modification of Policy and Utility Function in Rational Agents
    2016/05/10 by Tom Everitt, Daniel Filan, Everitt, Tom +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Reinforcement Learning in Robotics
  13. Path-Specific Objectives for Safer Agent Incentives
    2022/04/21 by Sebastian Farquhar, Ryan M. Carey, Farquhar, Sebastian +3 · 1 citation
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (stat.ML)
  14. Measuring Goal-Directedness
    2024/12/06 by Matt MacDermott, MacDermott, Matt, James Fox +5 · 2 citations
    Decision Sciences · #Evaluation and Performance Assessment