Adam Gleave
- Adversarial Policies Beat Superhuman Go AIs
2022/11/01 by Tony Tong Wang, Tony T. Wang, Adam Gleave +21 · 18 voices · 7 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Adversarial Policies: Attacking Deep Reinforcement Learning
2019/05/25 by Adam Gleave, Gleave, Adam, Michael D. Dennis +10 · 1 voice · 24 citations
Computer Science · #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics #Anomaly Detection Techniques and Applications
- Multi-Agent Risks from Advanced AI
2025/02/19 by Lewis Hammond, Hammond, Lewis, Alan Chan +89 · 3 voices · 37 citations
Social Sciences · #Ethics and Social Impacts of AI #cs.AI #cs.CY #cs.ET #cs.LG #cs.MA
- Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
2022/03/14 by Joar Skalse, Matthew Farrugia-Roberts, Skalse, Joar +7 · 5 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
- A Primer on Maximum Causal Entropy Inverse Reinforcement Learning
2022/03/22 by Adam Gleave, Sam Toyer, Gleave, Adam +1 · 4 citations
Computer Science · #Reinforcement Learning in Robotics #Evolutionary Algorithms and Applications
- Scaling Trends for Data Poisoning in LLMs
2024/08/06 by Dillon Bowen, Brendan Murphy, Bowen, Dillon +9 · 7 citations
Computer Science · Decision Sciences · #Privacy-Preserving Technologies in Data #Scientific Computing and Data Management
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Bengio, Yoshua, Tegan Maharaj +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Stable-Baselines3: Reliable Reinforcement Learning Implementations
2021/01/01 by Antonin Raffin, Ashley Hill, Adam Gleave +3 · 3 citations
Business, Management and Accounting · #Supply Chain and Inventory Management
- Understanding Learned Reward Functions
2020/12/10 by Eric J. Michaud, Michaud, Eric J., Adam Gleave +3 · 2 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
- Planning in a recurrent neural network that plays Sokoban
2024/07/22 by Mohammad Taufeeque, Philip Quirke, Taufeeque, Mohammad +11 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- Quantifying Differences in Reward Functions
2020/06/24 by Adam Gleave, Michael D. Dennis, Gleave, Adam +7 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Uncertainty Estimation for Language Reward Models
2022/03/14 by Adam Gleave, Geoffrey Irving, Gleave, Adam +1 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
2025/06/03 by Matthew Kowal, Kowal, Matthew, Jasper Timm +15 · 3 citations
Social Sciences · Computer Science · Psychology · #Misinformation and Its Impacts #Hate Speech and Cyberbullying Detection #Deception detection and forensic psychology
- Active Inverse Reward Design
2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- Preference Learning with Lie Detectors can Induce Honesty or Evasion
2025/05/20 by Chris Cundy, Adam Gleave, Cundy, Chris +1 · 3 citations
Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG)
- STARC: A General Framework For Quantifying Differences Between Reward Functions
2023/09/26 by Joar Skalse, Skalse, Joar, Lucy Farnik +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Large language models can effectively convince people to believe conspiracies
2026/01/08 by Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6 · 5 voices
#cs.AI #econ.GN #q-fin.EC