Gleave, Adam
- Adversarial Policies Beat Superhuman Go AIs
2022/11/01 by Tony T. Wang, Tony Tong Wang, Adam Gleave +21 · 18 voices · 7 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Adversarial Policies: Attacking Deep Reinforcement Learning
2019/05/25 by Adam Gleave, Michael Dennis, Gleave, Adam +10 · 1 voice · 25 citations
Computer Science · #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics #Anomaly Detection Techniques and Applications
- Multi-Agent Risks from Advanced AI
2025/02/19 by Lewis Hammond, Hammond, Lewis, Alan Chan +89 · 3 voices · 37 citations
Social Sciences · #Ethics and Social Impacts of AI #cs.AI #cs.CY #cs.ET #cs.LG #cs.MA
- Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
2022/03/14 by Joar Skalse, Matthew Farrugia-Roberts, Skalse, Joar +7 · 5 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
- imitation: Clean Imitation Learning Implementations
2022/11/22 by Gleave, Adam, Taufeeque, Mohammad, Rocamonde, Juan +7 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- A Primer on Maximum Causal Entropy Inverse Reinforcement Learning
2022/03/22 by Adam Gleave, Sam Toyer, Gleave, Adam +1 · 4 citations
Computer Science · #Reinforcement Learning in Robotics #Evolutionary Algorithms and Applications
- Scaling Trends for Data Poisoning in LLMs
2024/08/06 by Dillon Bowen, Bowen, Dillon, Brendan Murphy +9 · 7 citations
Computer Science · Decision Sciences · #Privacy-Preserving Technologies in Data #Scientific Computing and Data Management
- Scaling Trends in Language Model Robustness
2024/07/25 by Howe, Nikolaus, McKenzie, Ian, Hollinsworth, Oskar +5 · 7 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG)
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Exploiting Novel GPT-4 APIs
2023/12/21 by Pelrine, Kellin, Taufeeque, Mohammad, Zając, Michał +2 · 4 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG)
- Understanding Learned Reward Functions
2020/12/10 by Eric J. Michaud, Adam Gleave, Michaud, Eric J. +3 · 2 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
- On The Fragility of Learned Reward Functions
2023/01/09 by McKinney, Lev, Duan, Yawen, Krueger, David +1 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Planning in a recurrent neural network that plays Sokoban
2024/07/22 by Mohammad Taufeeque, Philip Quirke, Taufeeque, Mohammad +11 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- Quantifying Differences in Reward Functions
2020/06/24 by Adam Gleave, Michael D. Dennis, Gleave, Adam +7 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Uncertainty Estimation for Language Reward Models
2022/03/14 by Adam Gleave, Geoffrey Irving, Gleave, Adam +1 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Reducing Exploitability with Population Based Training
2022/08/10 by Czempin, Pavel, Gleave, Adam · 1 citation
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
2025/06/03 by Matthew Kowal, Jasper Timm, Kowal, Matthew +15 · 3 citations
Social Sciences · Computer Science · Psychology · #Misinformation and Its Impacts #Hate Speech and Cyberbullying Detection #Deception detection and forensic psychology
- Active Inverse Reward Design
2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- Preference Learning with Lie Detectors can Induce Honesty or Evasion
2025/05/20 by Chris Cundy, Cundy, Chris, Adam Gleave +1 · 3 citations
Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG)
- STARC: A General Framework For Quantifying Differences Between Reward Functions
2023/09/26 by Joar Skalse, Lucy Farnik, Skalse, Joar +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Can Go AIs be adversarially robust?
2024/06/18 by Tseng, Tom, McLean, Euan, Pelrine, Kellin +2 · 1 voice
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)