David Krueger
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
2025/01/28 by Jan Kulveit, Raymond Douglas, Kulveit, Jan +10 · 21 voices · 21 citations
Social Sciences · #Ethics and Social Impacts of AI
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 88 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
2020/04/15 by Miles Brundage, Shahar Avin, Brundage, Miles +115 · 2 voices · 21 citations
#cs.CY
- A Closer Look at Memorization in Deep Networks
2017/06/16 by Devansh Arpit, Arpit, Devansh, Stanisław Jastrzȩbski +19 · 71 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Visibility into AI Agents
2024/01/23 by Alan Chan, Chan, Alan, Carson Ezell +21 · 2 voices · 11 citations
Computer Science · Social Sciences · #Blockchain Technology Applications and Security #Cybercrime and Law Enforcement Studies #Ethics and Social Impacts of AI #cs.AI #cs.CY
- Scalable agent alignment via reward modeling: a research direction
2018/11/19 by Jan Leike, Leike, Jan, David Krueger +9 · 66 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- Out-of-Distribution Generalization via Risk Extrapolation (REx)
2020/03/02 by David Krueger, Krueger, David, Ethan Caballero +13 · 55 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Gaussian Processes and Bayesian Inference
- Defining and Characterizing Reward Hacking
2022/09/27 by Joar Skalse, Skalse, Joar, Nikolaus H. R. Howe +5 · 26 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
- Reward Model Ensembles Help Mitigate Overoptimization
2023/10/04 by Thomas Coste, Coste, Thomas, Robert Kirk +4 · 25 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Barez, Fazl, Tingchen Fu +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Towards Interpreting Visual Information Processing in Vision-Language Models
2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Active Reinforcement Learning: Observing Rewards at a Cost
2020/11/13 by David Krueger, Jan Leike, Krueger, David +5 · 5 citations
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
- Mechanistic Mode Connectivity
2022/11/15 by Ekdeep Singh Lubana, Lubana, Ekdeep Singh, Eric Bigelow +7 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Neural Networks and Applications
- AI Research Considerations for Human Existential Safety (ARCHES)
2020/05/30 by Andrew Critch, Critch, Andrew, David Krueger +1 · 4 citations
Computer Science · Social Sciences · #68T01 #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0 #Machine Learning (cs.LG)
- Hidden Incentives for Auto-Induced Distributional Shift
2020/09/19 by David Krueger, Krueger, David, Tegan Maharaj +3 · 4 citations
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
- Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
2024/11/02 by Luke Marks, Alasdair Paren, Marks, Luke +5 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
- Pitfalls of Evidence-Based AI Policy
2025/02/13 by Stephen Casper, Casper, Stephen, David Krueger +3 · 3 voices · 5 citations
#cs.CY
- Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
2024/10/22 by Itamar Pres, Laura Ruis, Pres, Itamar +5 · 7 citations
Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy
- Permissive Information-Flow Analysis for Large Language Models
2024/10/04 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Gaonkar, Radhika +16 · 7 citations
Computer Science · Decision Sciences · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024/10/11 by Tong Fu, Mrinank Sharma, Fu, Tingchen +9 · 7 citations
Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
- Stress-Testing Capability Elicitation With Password-Locked Models
2024/05/29 by Ryan Greenblatt, Fabien Roger, Greenblatt, Ryan +5 · 5 citations
Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
- Blockwise Self-Supervised Learning at Scale
2023/02/03 by Shoaib Ahmed Siddiqui, David Krueger, Siddiqui, Shoaib Ahmed +5 · 3 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Neural Networks and Reservoir Computing
- Influence Functions for Scalable Data Attribution in Diffusion Models
2024/10/17 by Bruno Mlodozeniec, Runa Eschenhagen, Mlodozeniec, Bruno +9 · 6 citations
Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Opinion Dynamics and Social Influence
- Implicit meta-learning may lead language models to trust more reliable sources
2023/10/23 by Dmitrii Krasheninnikov, Egor Krasheninnikov, Krasheninnikov, Dmitrii +7 · 2 voices · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Topic Modeling #cs.AI #cs.LG
- Affirmative safety: An approach to risk management for high-risk AI
2024/04/14 by Akash R. Wasil, Wasil, Akash R., Joshua Clymer +9 · 3 citations
Computer Science · Decision Sciences · Health Professions · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Occupational Health and Safety Research #Risk and Safety Analysis
- AI Researchers' Views on Automating AI R&D and Intelligence Explosions
2026/02/13 by Severin Field, Raymond Douglas, David Krueger · 1 voice
#cs.CY
- Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
2024/11/11 by Madeline Brumley, Brumley, Madeline, Kwon, Joe +5 · 3 citations
Psychology · #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Machine Learning (cs.LG)
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
2024/10/09 by Michael Lan, Philip Torr, Lan, Michael +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
2022/09/20 by Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Siddiqui, Shoaib Ahmed +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Thinker: Learning to Plan and Act
2023/07/27 by Stephen S. Chung, Ivan Anokhin, Chung, Stephen +3 · 1 citation
Computer Science · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation
- Fresh in memory: Training-order recency is linearly encoded in language model activations
2025/09/17 by Dmitrii Krasheninnikov, Richard E. Turner, Krasheninnikov, Dmitrii +3 · 3 voices · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Input Space Mode Connectivity in Deep Neural Networks
2024/09/09 by Jakub Vrábel, Ori Shem-Ur, Vrabel, Jakub +5 · 1 citation
Computer Science · Earth and Planetary Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Earthquake Detection and Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Neural Networks and Reservoir Computing #Seismology and Earthquake Studies #Statistical Mechanics (cond-mat.stat-mech)
- Learning to Forget using Hypernetworks
2024/12/01 by Jose Miguel Lara Rangel, Stefan Schoepf, Rangel, Jose Miguel Lara +6 · 1 citation
Social Sciences · #Educational Tools and Methods
- Understanding (Un)Reliability of Steering Vectors in Language Models
2025/05/28 by Joschka Braun, Carsten Eickhoff, Braun, Joschka +7 · 2 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling