vix.ing · top · new · best · stats · spec

Adam Gleave

  1. Adversarial Policies Beat Superhuman Go AIs
    2022/11/01 by Tony Tong Wang, Tony T. Wang, Adam Gleave +21 · 18 voices · 7 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
  2. Adversarial Policies: Attacking Deep Reinforcement Learning
    2019/05/25 by Adam Gleave, Gleave, Adam, Michael D. Dennis +10 · 1 voice · 24 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics #Anomaly Detection Techniques and Applications
  3. Multi-Agent Risks from Advanced AI
    2025/02/19 by Lewis Hammond, Hammond, Lewis, Alan Chan +89 · 3 voices · 37 citations
    Social Sciences · #Ethics and Social Impacts of AI #cs.AI #cs.CY #cs.ET #cs.LG #cs.MA
  4. Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
    2022/03/14 by Joar Skalse, Matthew Farrugia-Roberts, Skalse, Joar +7 · 5 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
  5. A Primer on Maximum Causal Entropy Inverse Reinforcement Learning
    2022/03/22 by Adam Gleave, Sam Toyer, Gleave, Adam +1 · 4 citations
    Computer Science · #Reinforcement Learning in Robotics #Evolutionary Algorithms and Applications
  6. Scaling Trends for Data Poisoning in LLMs
    2024/08/06 by Dillon Bowen, Brendan Murphy, Bowen, Dillon +9 · 7 citations
    Computer Science · Decision Sciences · #Privacy-Preserving Technologies in Data #Scientific Computing and Data Management
  7. The Singapore Consensus on Global AI Safety Research Priorities
    2025/06/25 by Yoshua Bengio, Bengio, Yoshua, Tegan Maharaj +171 · 2 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  8. Stable-Baselines3: Reliable Reinforcement Learning Implementations
    2021/01/01 by Antonin Raffin, Ashley Hill, Adam Gleave +3 · 3 citations
    Business, Management and Accounting · #Supply Chain and Inventory Management
  9. Understanding Learned Reward Functions
    2020/12/10 by Eric J. Michaud, Michaud, Eric J., Adam Gleave +3 · 2 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
  10. Planning in a recurrent neural network that plays Sokoban
    2024/07/22 by Mohammad Taufeeque, Philip Quirke, Taufeeque, Mohammad +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  11. Quantifying Differences in Reward Functions
    2020/06/24 by Adam Gleave, Michael D. Dennis, Gleave, Adam +7 · 1 citation
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  12. Uncertainty Estimation for Language Reward Models
    2022/03/14 by Adam Gleave, Geoffrey Irving, Gleave, Adam +1 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  13. It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
    2025/06/03 by Matthew Kowal, Kowal, Matthew, Jasper Timm +15 · 3 citations
    Social Sciences · Computer Science · Psychology · #Misinformation and Its Impacts #Hate Speech and Cyberbullying Detection #Deception detection and forensic psychology
  14. Active Inverse Reward Design
    2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  15. Preference Learning with Lie Detectors can Induce Honesty or Evasion
    2025/05/20 by Chris Cundy, Adam Gleave, Cundy, Chris +1 · 3 citations
    Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG)
  16. STARC: A General Framework For Quantifying Differences Between Reward Functions
    2023/09/26 by Joar Skalse, Skalse, Joar, Lucy Farnik +9 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
  17. STACK: Adversarial Attacks on LLM Safeguard Pipelines
    2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
  18. Large language models can effectively convince people to believe conspiracies
    2026/01/08 by Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6 · 5 voices
    #cs.AI #econ.GN #q-fin.EC