vix.ing · top · new · best · stats · spec

Mindermann, Sören

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 104 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. The Alignment Problem from a Deep Learning Perspective
    2022/08/30 by Ngo, Richard, Chan, Lawrence, Mindermann, Sören · 8 voices · 40 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Alignment faking in large language models
    2024/12/18 by Ryan Greenblatt, Carson Denison, Greenblatt, Ryan +38 · 16 voices · 76 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt
    2022/06/14 by Sören Mindermann, Mindermann, Sören, Jan Brauner +19 · 28 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  5. Open Problems in Machine Unlearning for AI Safety
    2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 12 citations
    Engineering · #Fault Detection and Control Systems
  6. How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
    2023/09/26 by Lorenzo Pacchiardi, Pacchiardi, Lorenzo, Alex Chan +13 · 18 citations
    Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  7. Scientist AI Needs a Government: Structural Governance as the Missing Layer in Non-Agentic AI Safety
    2025/02/21 by Bengio, Yoshua, Michael K. Cohen, Damiano Fornasiere +22 · 25 citations
    Physics and Astronomy · #Space Science and Extraterrestrial Life
  8. International AI Safety Report
    2025/01/29 by Yoshua Bengio, Sören Mindermann, Bengio, Yoshua +180 · 20 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Identifying Causal-Effect Inference Failure with Uncertainty-Aware Models
    2020/07/01 by Jesson, Andrew, Mindermann, Sören, Shalit, Uri +1 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  10. Specific versus General Principles for Constitutional AI
    2023/10/20 by Sandipan Kundu, Kundu, Sandipan, Yuntao Bai +69 · 5 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
  11. The Singapore Consensus on Global AI Safety Research Priorities
    2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  12. International Scientific Report on the Safety of Advanced AI (Interim Report)
    2024/11/05 by Bengio, Yoshua, Mindermann, Sören, Privitera, Daniel +41 · 4 citations
    #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
  13. Active Inverse Reward Design
    2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  14. In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
    2025/04/17 by Ben Bucknall, Saad Siddiqui, Bucknall, Ben +41 · 2 voices · 2 citations
    #cs.CY
  15. Occam's razor is insufficient to infer the preferences of irrational agents
    2017/12/15 by Armstrong, Stuart, Mindermann, Sören · 1 citation
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
  16. Bare Minimum Mitigations for Autonomous AI Development
    2025/04/21 by Clymer, Joshua, Duan, Isabella, Cundy, Chris +10 · 1 citation
    #Computers and Society (cs.CY) #FOS: Computer and information sciences
  17. International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
    2025/10/15 by Yoshua Bengio, Bengio, Yoshua, Stephen Clare +134 · 1 citation
    Social Sciences · #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences