vix.ing · top · new · best · stats · spec

Gleave, Adam

  1. Adversarial Policies Beat Superhuman Go AIs
    2022/11/01 by Tony T. Wang, Tony Tong Wang, Adam Gleave +21 · 18 voices · 7 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
  2. Adversarial Policies: Attacking Deep Reinforcement Learning
    2019/05/25 by Adam Gleave, Michael Dennis, Gleave, Adam +10 · 1 voice · 25 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics #Anomaly Detection Techniques and Applications
  3. Multi-Agent Risks from Advanced AI
    2025/02/19 by Lewis Hammond, Hammond, Lewis, Alan Chan +89 · 3 voices · 37 citations
    Social Sciences · #Ethics and Social Impacts of AI #cs.AI #cs.CY #cs.ET #cs.LG #cs.MA
  4. Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
    2022/03/14 by Joar Skalse, Matthew Farrugia-Roberts, Skalse, Joar +7 · 5 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
  5. imitation: Clean Imitation Learning Implementations
    2022/11/22 by Gleave, Adam, Taufeeque, Mohammad, Rocamonde, Juan +7 · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. A Primer on Maximum Causal Entropy Inverse Reinforcement Learning
    2022/03/22 by Adam Gleave, Sam Toyer, Gleave, Adam +1 · 4 citations
    Computer Science · #Reinforcement Learning in Robotics #Evolutionary Algorithms and Applications
  7. Scaling Trends for Data Poisoning in LLMs
    2024/08/06 by Dillon Bowen, Bowen, Dillon, Brendan Murphy +9 · 7 citations
    Computer Science · Decision Sciences · #Privacy-Preserving Technologies in Data #Scientific Computing and Data Management
  8. Scaling Trends in Language Model Robustness
    2024/07/25 by Howe, Nikolaus, McKenzie, Ian, Hollinsworth, Oskar +5 · 7 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG)
  9. The Singapore Consensus on Global AI Safety Research Priorities
    2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  10. Exploiting Novel GPT-4 APIs
    2023/12/21 by Pelrine, Kellin, Taufeeque, Mohammad, Zając, Michał +2 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG)
  11. Understanding Learned Reward Functions
    2020/12/10 by Eric J. Michaud, Adam Gleave, Michaud, Eric J. +3 · 2 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
  12. On The Fragility of Learned Reward Functions
    2023/01/09 by McKinney, Lev, Duan, Yawen, Krueger, David +1 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. Planning in a recurrent neural network that plays Sokoban
    2024/07/22 by Mohammad Taufeeque, Philip Quirke, Taufeeque, Mohammad +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  14. Quantifying Differences in Reward Functions
    2020/06/24 by Adam Gleave, Michael D. Dennis, Gleave, Adam +7 · 1 citation
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  15. Uncertainty Estimation for Language Reward Models
    2022/03/14 by Adam Gleave, Geoffrey Irving, Gleave, Adam +1 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  16. Reducing Exploitability with Population Based Training
    2022/08/10 by Czempin, Pavel, Gleave, Adam · 1 citation
    #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  17. It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
    2025/06/03 by Matthew Kowal, Jasper Timm, Kowal, Matthew +15 · 3 citations
    Social Sciences · Computer Science · Psychology · #Misinformation and Its Impacts #Hate Speech and Cyberbullying Detection #Deception detection and forensic psychology
  18. Active Inverse Reward Design
    2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  19. Preference Learning with Lie Detectors can Induce Honesty or Evasion
    2025/05/20 by Chris Cundy, Cundy, Chris, Adam Gleave +1 · 3 citations
    Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG)
  20. STARC: A General Framework For Quantifying Differences Between Reward Functions
    2023/09/26 by Joar Skalse, Lucy Farnik, Skalse, Joar +9 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
  21. STACK: Adversarial Attacks on LLM Safeguard Pipelines
    2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
  22. Can Go AIs be adversarially robust?
    2024/06/18 by Tseng, Tom, McLean, Euan, Pelrine, Kellin +2 · 1 voice
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)