Sören Mindermann
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 64 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Inferring the effectiveness of government interventions against COVID-19
2020/12/16 by Jan Brauner, Sören Mindermann, Mrinank Sharma +16 · 14 citations
Economics, Econometrics and Finance · Mathematics · Psychology · #COVID-19 Pandemic Impacts #COVID-19 and Mental Health #COVID-19 epidemiological studies
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
2023/09/26 by Lorenzo Pacchiardi, Alex Chan, Pacchiardi, Lorenzo +13 · 13 citations
Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Scientist AI Needs a Government: Structural Governance as the Missing Layer in Non-Agentic AI Safety
2025/02/21 by Bengio, Yoshua, Michael K. Cohen, Damiano Fornasiere +22 · 19 citations
Physics and Astronomy · #Space Science and Extraterrestrial Life
- International AI Safety Report
2025/01/29 by Yoshua Bengio, Sören Mindermann, Bengio, Yoshua +180 · 17 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Specific versus General Principles for Constitutional AI
2023/10/20 by Sandipan Kundu, Kundu, Sandipan, Yuntao Bai +69 · 4 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Agentic Misalignment: How LLMs Could Be Insider Threats
2025/10/05 by Aengus Lynch, Benjamin Wright, Lynch, Aengus +12 · 20 citations
Business, Management and Accounting · Computer Science · #Securities Regulation and Market Practices #Corporate Insolvency and Governance #Cybercrime and Law Enforcement Studies
- Understanding the effectiveness of government interventions against the resurgence of COVID-19 in Europe
2021/10/05 by Mrinank Sharma, Sören Mindermann, Darren Smith +21 · 1 voice · 2 citations
Economics, Econometrics and Finance · Mathematics · Social Sciences · #COVID-19 Pandemic Impacts #COVID-19 epidemiological studies #Vaccine Coverage and Hesitancy
- Active Inverse Reward Design
2018/09/09 by Sören Mindermann, Rohin Shah, Mindermann, Sören +5 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
2025/04/17 by Ben Bucknall, Saad Siddiqui, Bucknall, Ben +41 · 2 voices · 1 citation
#cs.CY
- International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
2025/10/15 by Yoshua Bengio, Stephen Clare, Bengio, Yoshua +134 · 1 citation
Social Sciences · #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences