vix.ing · top · new · best · stats · spec

David Krueger

  1. Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
    2025/01/28 by Jan Kulveit, Raymond Douglas, Kulveit, Jan +10 · 21 voices · 21 citations
    Social Sciences · #Ethics and Social Impacts of AI
  2. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 88 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  3. Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
    2020/04/15 by Miles Brundage, Shahar Avin, Brundage, Miles +115 · 2 voices · 21 citations
    #cs.CY
  4. A Closer Look at Memorization in Deep Networks
    2017/06/16 by Devansh Arpit, Arpit, Devansh, Stanisław Jastrzȩbski +19 · 71 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  5. Visibility into AI Agents
    2024/01/23 by Alan Chan, Chan, Alan, Carson Ezell +21 · 2 voices · 11 citations
    Computer Science · Social Sciences · #Blockchain Technology Applications and Security #Cybercrime and Law Enforcement Studies #Ethics and Social Impacts of AI #cs.AI #cs.CY
  6. Scalable agent alignment via reward modeling: a research direction
    2018/11/19 by Jan Leike, Leike, Jan, David Krueger +9 · 66 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  7. Out-of-Distribution Generalization via Risk Extrapolation (REx)
    2020/03/02 by David Krueger, Krueger, David, Ethan Caballero +13 · 55 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Gaussian Processes and Bayesian Inference
  8. Defining and Characterizing Reward Hacking
    2022/09/27 by Joar Skalse, Skalse, Joar, Nikolaus H. R. Howe +5 · 26 citations
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
  9. Reward Model Ensembles Help Mitigate Overoptimization
    2023/10/04 by Thomas Coste, Coste, Thomas, Robert Kirk +4 · 25 citations
    Computer Science · #Machine Learning and Data Classification #Topic Modeling
  10. Open Problems in Machine Unlearning for AI Safety
    2025/01/09 by Fazl Barez, Barez, Fazl, Tingchen Fu +39 · 6 voices · 9 citations
    Engineering · #Fault Detection and Control Systems
  11. Towards Interpreting Visual Information Processing in Vision-Language Models
    2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  12. Active Reinforcement Learning: Observing Rewards at a Cost
    2020/11/13 by David Krueger, Jan Leike, Krueger, David +5 · 5 citations
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
  13. Mechanistic Mode Connectivity
    2022/11/15 by Ekdeep Singh Lubana, Lubana, Ekdeep Singh, Eric Bigelow +7 · 6 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Neural Networks and Applications
  14. AI Research Considerations for Human Existential Safety (ARCHES)
    2020/05/30 by Andrew Critch, Critch, Andrew, David Krueger +1 · 4 citations
    Computer Science · Social Sciences · #68T01 #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0 #Machine Learning (cs.LG)
  15. Hidden Incentives for Auto-Induced Distributional Shift
    2020/09/19 by David Krueger, Krueger, David, Tegan Maharaj +3 · 4 citations
    Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
  16. Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
    2024/11/02 by Luke Marks, Alasdair Paren, Marks, Luke +5 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
  17. Pitfalls of Evidence-Based AI Policy
    2025/02/13 by Stephen Casper, Casper, Stephen, David Krueger +3 · 3 voices · 5 citations
    #cs.CY
  18. Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
    2024/10/22 by Itamar Pres, Laura Ruis, Pres, Itamar +5 · 7 citations
    Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy
  19. Permissive Information-Flow Analysis for Large Language Models
    2024/10/04 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Gaonkar, Radhika +16 · 7 citations
    Computer Science · Decision Sciences · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  20. PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
    2024/10/11 by Tong Fu, Mrinank Sharma, Fu, Tingchen +9 · 7 citations
    Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
  21. Stress-Testing Capability Elicitation With Password-Locked Models
    2024/05/29 by Ryan Greenblatt, Fabien Roger, Greenblatt, Ryan +5 · 5 citations
    Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
  22. Blockwise Self-Supervised Learning at Scale
    2023/02/03 by Shoaib Ahmed Siddiqui, David Krueger, Siddiqui, Shoaib Ahmed +5 · 3 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Neural Networks and Reservoir Computing
  23. Influence Functions for Scalable Data Attribution in Diffusion Models
    2024/10/17 by Bruno Mlodozeniec, Runa Eschenhagen, Mlodozeniec, Bruno +9 · 6 citations
    Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Opinion Dynamics and Social Influence
  24. Implicit meta-learning may lead language models to trust more reliable sources
    2023/10/23 by Dmitrii Krasheninnikov, Egor Krasheninnikov, Krasheninnikov, Dmitrii +7 · 2 voices · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Topic Modeling #cs.AI #cs.LG
  25. Affirmative safety: An approach to risk management for high-risk AI
    2024/04/14 by Akash R. Wasil, Wasil, Akash R., Joshua Clymer +9 · 3 citations
    Computer Science · Decision Sciences · Health Professions · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Occupational Health and Safety Research #Risk and Safety Analysis
  26. AI Researchers' Views on Automating AI R&D and Intelligence Explosions
    2026/02/13 by Severin Field, Raymond Douglas, David Krueger · 1 voice
    #cs.CY
  27. Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
    2024/11/11 by Madeline Brumley, Brumley, Madeline, Kwon, Joe +5 · 3 citations
    Psychology · #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Machine Learning (cs.LG)
  28. Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
    2024/10/09 by Michael Lan, Philip Torr, Lan, Michael +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  29. Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
    2022/09/20 by Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Siddiqui, Shoaib Ahmed +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Machine Learning and Data Classification
  30. Thinker: Learning to Plan and Act
    2023/07/27 by Stephen S. Chung, Ivan Anokhin, Chung, Stephen +3 · 1 citation
    Computer Science · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation
  31. Fresh in memory: Training-order recency is linearly encoded in language model activations
    2025/09/17 by Dmitrii Krasheninnikov, Richard E. Turner, Krasheninnikov, Dmitrii +3 · 3 voices · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  32. Input Space Mode Connectivity in Deep Neural Networks
    2024/09/09 by Jakub Vrábel, Ori Shem-Ur, Vrabel, Jakub +5 · 1 citation
    Computer Science · Earth and Planetary Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Earthquake Detection and Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Neural Networks and Reservoir Computing #Seismology and Earthquake Studies #Statistical Mechanics (cond-mat.stat-mech)
  33. Learning to Forget using Hypernetworks
    2024/12/01 by Jose Miguel Lara Rangel, Stefan Schoepf, Rangel, Jose Miguel Lara +6 · 1 citation
    Social Sciences · #Educational Tools and Methods
  34. Understanding (Un)Reliability of Steering Vectors in Language Models
    2025/05/28 by Joschka Braun, Carsten Eickhoff, Braun, Joschka +7 · 2 citations
    Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling