vix.ing · top · new · best · stats · spec

Krueger, David

  1. Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
    2025/01/28 by Jan Kulveit, Raymond Douglas, Kulveit, Jan +10 · 21 voices · 21 citations
    Social Sciences · #Ethics and Social Impacts of AI
  2. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 87 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  3. NICE: Non-linear Independent Components Estimation
    2014/10/30 by Dinh, Laurent, Krueger, David, Bengio, Yoshua · 73 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  4. Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
    2020/04/15 by Miles Brundage, Brundage, Miles, Shahar Avin +115 · 2 voices · 21 citations
    #cs.CY
  5. Visibility into AI Agents
    2024/01/23 by Alan Chan, Chan, Alan, Carson Ezell +21 · 2 voices · 11 citations
    Computer Science · Social Sciences · #Blockchain Technology Applications and Security #Cybercrime and Law Enforcement Studies #Ethics and Social Impacts of AI #cs.AI #cs.CY
  6. A Closer Look at Memorization in Deep Networks
    2017/06/16 by Devansh Arpit, Stanisław Jastrzȩbski, Arpit, Devansh +19 · 69 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  7. Scalable agent alignment via reward modeling: a research direction
    2018/11/19 by Jan Leike, David Krueger, Leike, Jan +9 · 64 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  8. Out-of-Distribution Generalization via Risk Extrapolation (REx)
    2020/03/02 by David Krueger, Ethan Caballero, Krueger, David +13 · 55 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Gaussian Processes and Bayesian Inference
  9. Defining and Characterizing Reward Hacking
    2022/09/27 by Joar Skalse, Skalse, Joar, Nikolaus H. R. Howe +5 · 26 citations
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
  10. Neural Autoregressive Flows
    2018/04/03 by Huang, Chin-Wei, Krueger, David, Lacoste, Alexandre +1 · 14 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  11. Reward Model Ensembles Help Mitigate Overoptimization
    2023/10/04 by Thomas Coste, Coste, Thomas, Anwar, Usman +4 · 25 citations
    Computer Science · #Machine Learning and Data Classification #Topic Modeling
  12. Open Problems in Machine Unlearning for AI Safety
    2025/01/09 by Fazl Barez, Barez, Fazl, Tingchen Fu +39 · 6 voices · 9 citations
    Engineering · #Fault Detection and Control Systems
  13. Foundational Challenges in Assuring Alignment and Safety of Large Language Models
    2024/04/15 by Anwar, Usman, Saparov, Abulhair, Rando, Javier +39 · 22 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  14. Goal Misgeneralization in Deep Reinforcement Learning
    2021/05/28 by Langosco, Lauro, Koch, Jack, Sharkey, Lee +3 · 10 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  15. Towards Interpreting Visual Information Processing in Vision-Language Models
    2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  16. Safety Cases: How to Justify the Safety of Advanced AI Systems
    2024/03/15 by Clymer, Joshua, Gabrieli, Nick, Krueger, David +1 · 12 citations
    #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
  17. Bayesian Hypernetworks
    2017/10/13 by Krueger, David, Huang, Chin-Wei, Islam, Riashat +3 · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  18. Characterizing Manipulation from AI Systems
    2023/03/16 by Carroll, Micah, Chan, Alan, Ashton, Henry +1 · 7 citations
    #Computers and Society (cs.CY) #FOS: Computer and information sciences
  19. Active Reinforcement Learning: Observing Rewards at a Cost
    2020/11/13 by David Krueger, Krueger, David, Jan Leike +5 · 5 citations
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
  20. Mechanistic Mode Connectivity
    2022/11/15 by Ekdeep Singh Lubana, Lubana, Ekdeep Singh, Eric Bigelow +7 · 6 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Neural Networks and Applications
  21. Unifying Grokking and Double Descent
    2023/03/10 by Davies, Xander, Langosco, Lauro, Krueger, David · 6 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  22. AI Research Considerations for Human Existential Safety (ARCHES)
    2020/05/30 by Andrew Critch, David Krueger, Critch, Andrew +1 · 4 citations
    Computer Science · Social Sciences · #68T01 #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0 #Machine Learning (cs.LG)
  23. Hidden Incentives for Auto-Induced Distributional Shift
    2020/09/19 by David Krueger, Tegan Maharaj, Krueger, David +3 · 4 citations
    Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
  24. Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
    2024/11/02 by Luke Marks, Marks, Luke, Alasdair Paren +5 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
  25. Pitfalls of Evidence-Based AI Policy
    2025/02/13 by Stephen Casper, Casper, Stephen, David Krueger +3 · 3 voices · 5 citations
    #cs.CY
  26. Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
    2024/10/22 by Itamar Pres, Laura Ruis, Pres, Itamar +5 · 7 citations
    Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy
  27. Permissive Information-Flow Analysis for Large Language Models
    2024/10/04 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Boris Köpf +16 · 7 citations
    Computer Science · Decision Sciences · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  28. PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
    2024/10/11 by Tong Fu, Mrinank Sharma, Fu, Tingchen +9 · 7 citations
    Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
  29. Stress-Testing Capability Elicitation With Password-Locked Models
    2024/05/29 by Ryan Greenblatt, Fabien Roger, Greenblatt, Ryan +5 · 5 citations
    Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
  30. Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
    2025/02/27 by Wehner, Jan, Abdelnabi, Sahar, Tan, Daniel +2 · 7 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  31. Blockwise Self-Supervised Learning at Scale
    2023/02/03 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, David Krueger +5 · 3 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Neural Networks and Reservoir Computing
  32. Influence Functions for Scalable Data Attribution in Diffusion Models
    2024/10/17 by Bruno Mlodozeniec, Runa Eschenhagen, Mlodozeniec, Bruno +9 · 6 citations
    Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Opinion Dynamics and Social Influence
  33. A deeper look at depth pruning of LLMs
    2024/07/23 by Siddiqui, Shoaib Ahmed, Dong, Xin, Heinrich, Greg +4 · 4 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  34. Implicit meta-learning may lead language models to trust more reliable sources
    2023/10/23 by Dmitrii Krasheninnikov, Krasheninnikov, Dmitrii, Egor Krasheninnikov +7 · 2 voices · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Topic Modeling #cs.AI #cs.LG
  35. Detecting High-Stakes Interactions with Activation Probes
    2025/06/12 by McKenzie, Alex, Pawar, Urja, Blandfort, Phil +4 · 10 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  36. Affirmative safety: An approach to risk management for high-risk AI
    2024/04/14 by Akash R. Wasil, Wasil, Akash R., Joshua Clymer +9 · 3 citations
    Computer Science · Decision Sciences · Health Professions · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Occupational Health and Safety Research #Risk and Safety Analysis
  37. Analyzing (In)Abilities of SAEs via Formal Languages
    2024/10/15 by Menon, Abhinav, Shrivastava, Manish, Krueger, David +1 · 4 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  38. IDs for AI Systems
    2024/06/17 by Chan, Alan, Kolt, Noam, Wills, Peter +7 · 3 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
  39. On The Fragility of Learned Reward Functions
    2023/01/09 by McKinney, Lev, Duan, Yawen, Krueger, David +1 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  40. Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
    2016/06/03 by Krueger, David, Maharaj, Tegan, Kramár, János +7 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
  41. Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
    2024/11/11 by Madeline Brumley, Brumley, Madeline, David Krueger +5 · 3 citations
    Psychology · #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Machine Learning (cs.LG)
  42. A Generative Model of Symmetry Transformations
    2024/03/04 by Allingham, James Urquhart, Mlodozeniec, Bruno Kacper, Padhy, Shreyas +5 · 2 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  43. Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
    2024/10/09 by Michael Lan, Lan, Michael, Philip Torr +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  44. Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
    2022/09/20 by Shoaib Ahmed Siddiqui, Siddiqui, Shoaib Ahmed, Nitarshan Rajkumar +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Machine Learning and Data Classification
  45. Interpreting Emergent Planning in Model-Free Reinforcement Learning
    2025/04/02 by Bush, Thomas, Chung, Stephen, Anwar, Usman +2 · 3 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  46. Exploring the design space of deep-learning-based weather forecasting systems
    2024/10/09 by Siddiqui, Shoaib Ahmed, Kossaifi, Jean, Bonev, Boris +4 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  47. Thinker: Learning to Plan and Act
    2023/07/27 by Stephen S. Chung, Chung, Stephen, Ivan Anokhin +3 · 1 citation
    Computer Science · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation
  48. Towards Out-of-Distribution Adversarial Robustness
    2022/10/06 by Ibrahim, Adam, Guille-Escuret, Charles, Mitliagkas, Ioannis +3 · 1 citation
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  49. Fresh in memory: Training-order recency is linearly encoded in language model activations
    2025/09/17 by Dmitrii Krasheninnikov, Krasheninnikov, Dmitrii, Richard E. Turner +3 · 3 voices · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  50. Interpreting Learned Feedback Patterns in Large Language Models
    2023/10/12 by Luke Marks, Amir Abdullah, Marks, Luke +9 · 1 citation
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  51. Input Space Mode Connectivity in Deep Neural Networks
    2024/09/09 by Jakub Vrábel, Vrabel, Jakub, Ori Shem-Ur +5 · 1 citation
    Computer Science · Earth and Planetary Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Earthquake Detection and Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Neural Networks and Reservoir Computing #Seismology and Earthquake Studies #Statistical Mechanics (cond-mat.stat-mech)
  52. Distributional Training Data Attribution: What do Influence Functions Sample?
    2025/06/15 by Mlodozeniec, Bruno, Reid, Isaac, Power, Sam +4 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  53. Learning to Forget using Hypernetworks
    2024/12/01 by Jose Miguel Lara Rangel, Rangel, Jose Miguel Lara, Stefan Schoepf +6 · 1 citation
    Social Sciences · #Educational Tools and Methods
  54. Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
    2024/11/07 by Anwar, Usman, Von Oswald, Johannes, Kirsch, Louis +2 · 1 citation
    #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  55. Understanding (Un)Reliability of Steering Vectors in Language Models
    2025/05/28 by Joschka Braun, Braun, Joschka, Carsten Eickhoff +7 · 2 citations
    Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling