Krueger, David
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
2025/01/28 by Jan Kulveit, Raymond Douglas, Kulveit, Jan +10 · 21 voices · 21 citations
Social Sciences · #Ethics and Social Impacts of AI
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 87 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- NICE: Non-linear Independent Components Estimation
2014/10/30 by Dinh, Laurent, Krueger, David, Bengio, Yoshua · 73 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
2020/04/15 by Miles Brundage, Brundage, Miles, Shahar Avin +115 · 2 voices · 21 citations
#cs.CY
- Visibility into AI Agents
2024/01/23 by Alan Chan, Chan, Alan, Carson Ezell +21 · 2 voices · 11 citations
Computer Science · Social Sciences · #Blockchain Technology Applications and Security #Cybercrime and Law Enforcement Studies #Ethics and Social Impacts of AI #cs.AI #cs.CY
- A Closer Look at Memorization in Deep Networks
2017/06/16 by Devansh Arpit, Stanisław Jastrzȩbski, Arpit, Devansh +19 · 69 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Scalable agent alignment via reward modeling: a research direction
2018/11/19 by Jan Leike, David Krueger, Leike, Jan +9 · 64 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- Out-of-Distribution Generalization via Risk Extrapolation (REx)
2020/03/02 by David Krueger, Ethan Caballero, Krueger, David +13 · 55 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Gaussian Processes and Bayesian Inference
- Defining and Characterizing Reward Hacking
2022/09/27 by Joar Skalse, Skalse, Joar, Nikolaus H. R. Howe +5 · 26 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
- Neural Autoregressive Flows
2018/04/03 by Huang, Chin-Wei, Krueger, David, Lacoste, Alexandre +1 · 14 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Reward Model Ensembles Help Mitigate Overoptimization
2023/10/04 by Thomas Coste, Coste, Thomas, Anwar, Usman +4 · 25 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Barez, Fazl, Tingchen Fu +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
2024/04/15 by Anwar, Usman, Saparov, Abulhair, Rando, Javier +39 · 22 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Goal Misgeneralization in Deep Reinforcement Learning
2021/05/28 by Langosco, Lauro, Koch, Jack, Sharkey, Lee +3 · 10 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Towards Interpreting Visual Information Processing in Vision-Language Models
2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Safety Cases: How to Justify the Safety of Advanced AI Systems
2024/03/15 by Clymer, Joshua, Gabrieli, Nick, Krueger, David +1 · 12 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
- Bayesian Hypernetworks
2017/10/13 by Krueger, David, Huang, Chin-Wei, Islam, Riashat +3 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Characterizing Manipulation from AI Systems
2023/03/16 by Carroll, Micah, Chan, Alan, Ashton, Henry +1 · 7 citations
#Computers and Society (cs.CY) #FOS: Computer and information sciences
- Active Reinforcement Learning: Observing Rewards at a Cost
2020/11/13 by David Krueger, Krueger, David, Jan Leike +5 · 5 citations
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
- Mechanistic Mode Connectivity
2022/11/15 by Ekdeep Singh Lubana, Lubana, Ekdeep Singh, Eric Bigelow +7 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Neural Networks and Applications
- Unifying Grokking and Double Descent
2023/03/10 by Davies, Xander, Langosco, Lauro, Krueger, David · 6 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- AI Research Considerations for Human Existential Safety (ARCHES)
2020/05/30 by Andrew Critch, David Krueger, Critch, Andrew +1 · 4 citations
Computer Science · Social Sciences · #68T01 #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0 #Machine Learning (cs.LG)
- Hidden Incentives for Auto-Induced Distributional Shift
2020/09/19 by David Krueger, Tegan Maharaj, Krueger, David +3 · 4 citations
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
- Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
2024/11/02 by Luke Marks, Marks, Luke, Alasdair Paren +5 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
- Pitfalls of Evidence-Based AI Policy
2025/02/13 by Stephen Casper, Casper, Stephen, David Krueger +3 · 3 voices · 5 citations
#cs.CY
- Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
2024/10/22 by Itamar Pres, Laura Ruis, Pres, Itamar +5 · 7 citations
Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy
- Permissive Information-Flow Analysis for Large Language Models
2024/10/04 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Boris Köpf +16 · 7 citations
Computer Science · Decision Sciences · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024/10/11 by Tong Fu, Mrinank Sharma, Fu, Tingchen +9 · 7 citations
Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
- Stress-Testing Capability Elicitation With Password-Locked Models
2024/05/29 by Ryan Greenblatt, Fabien Roger, Greenblatt, Ryan +5 · 5 citations
Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
- Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
2025/02/27 by Wehner, Jan, Abdelnabi, Sahar, Tan, Daniel +2 · 7 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Blockwise Self-Supervised Learning at Scale
2023/02/03 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, David Krueger +5 · 3 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Neural Networks and Reservoir Computing
- Influence Functions for Scalable Data Attribution in Diffusion Models
2024/10/17 by Bruno Mlodozeniec, Runa Eschenhagen, Mlodozeniec, Bruno +9 · 6 citations
Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Opinion Dynamics and Social Influence
- A deeper look at depth pruning of LLMs
2024/07/23 by Siddiqui, Shoaib Ahmed, Dong, Xin, Heinrich, Greg +4 · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Implicit meta-learning may lead language models to trust more reliable sources
2023/10/23 by Dmitrii Krasheninnikov, Krasheninnikov, Dmitrii, Egor Krasheninnikov +7 · 2 voices · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Topic Modeling #cs.AI #cs.LG
- Detecting High-Stakes Interactions with Activation Probes
2025/06/12 by McKenzie, Alex, Pawar, Urja, Blandfort, Phil +4 · 10 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Affirmative safety: An approach to risk management for high-risk AI
2024/04/14 by Akash R. Wasil, Wasil, Akash R., Joshua Clymer +9 · 3 citations
Computer Science · Decision Sciences · Health Professions · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Occupational Health and Safety Research #Risk and Safety Analysis
- Analyzing (In)Abilities of SAEs via Formal Languages
2024/10/15 by Menon, Abhinav, Shrivastava, Manish, Krueger, David +1 · 4 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- IDs for AI Systems
2024/06/17 by Chan, Alan, Kolt, Noam, Wills, Peter +7 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- On The Fragility of Learned Reward Functions
2023/01/09 by McKinney, Lev, Duan, Yawen, Krueger, David +1 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
2016/06/03 by Krueger, David, Maharaj, Tegan, Kramár, János +7 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
2024/11/11 by Madeline Brumley, Brumley, Madeline, David Krueger +5 · 3 citations
Psychology · #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Machine Learning (cs.LG)
- A Generative Model of Symmetry Transformations
2024/03/04 by Allingham, James Urquhart, Mlodozeniec, Bruno Kacper, Padhy, Shreyas +5 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
2024/10/09 by Michael Lan, Lan, Michael, Philip Torr +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
2022/09/20 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Nitarshan Rajkumar +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Interpreting Emergent Planning in Model-Free Reinforcement Learning
2025/04/02 by Bush, Thomas, Chung, Stephen, Anwar, Usman +2 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Exploring the design space of deep-learning-based weather forecasting systems
2024/10/09 by Siddiqui, Shoaib Ahmed, Kossaifi, Jean, Bonev, Boris +4 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Thinker: Learning to Plan and Act
2023/07/27 by Stephen S. Chung, Chung, Stephen, Ivan Anokhin +3 · 1 citation
Computer Science · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation
- Towards Out-of-Distribution Adversarial Robustness
2022/10/06 by Ibrahim, Adam, Guille-Escuret, Charles, Mitliagkas, Ioannis +3 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Fresh in memory: Training-order recency is linearly encoded in language model activations
2025/09/17 by Dmitrii Krasheninnikov, Krasheninnikov, Dmitrii, Richard E. Turner +3 · 3 voices · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Interpreting Learned Feedback Patterns in Large Language Models
2023/10/12 by Luke Marks, Amir Abdullah, Marks, Luke +9 · 1 citation
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Input Space Mode Connectivity in Deep Neural Networks
2024/09/09 by Jakub Vrábel, Vrabel, Jakub, Ori Shem-Ur +5 · 1 citation
Computer Science · Earth and Planetary Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Earthquake Detection and Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Neural Networks and Reservoir Computing #Seismology and Earthquake Studies #Statistical Mechanics (cond-mat.stat-mech)
- Distributional Training Data Attribution: What do Influence Functions Sample?
2025/06/15 by Mlodozeniec, Bruno, Reid, Isaac, Power, Sam +4 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Learning to Forget using Hypernetworks
2024/12/01 by Jose Miguel Lara Rangel, Rangel, Jose Miguel Lara, Stefan Schoepf +6 · 1 citation
Social Sciences · #Educational Tools and Methods
- Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
2024/11/07 by Anwar, Usman, Von Oswald, Johannes, Kirsch, Louis +2 · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Understanding (Un)Reliability of Steering Vectors in Language Models
2025/05/28 by Joschka Braun, Braun, Joschka, Carsten Eickhoff +7 · 2 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling