vix.ing · top · new · best · stats · spec

Hubinger, Evan

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Korbak, Tomek, Mikita Balesni +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  3. Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
    2023/08/07 by Roger Grosse, Juhan Bae, Grosse, Roger +31 · 3 voices · 49 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
  4. Alignment faking in large language models
    2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 64 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  5. Measuring Faithfulness in Chain-of-Thought Reasoning
    2023/07/17 by Tamera Lanham, Lanham, Tamera, Anna Chen +59 · 4 voices · 83 citations
    Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
  6. Steering Llama 2 via Contrastive Activation Addition
    2023/12/09 by Panickssery, Nina, Nick Gabrieli, Julian Schulz +8 · 159 citations
    Computer Science · #Interactive and Immersive Displays
  7. Discovering Language Model Behaviors with Model-Written Evaluations
    2022/12/19 by Ethan Perez, Perez, Ethan, Sam Ringer +123 · 120 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
  8. Risks from Learned Optimization in Advanced Machine Learning Systems
    2019/06/05 by Evan Hubinger, Chris van Merwijk, Hubinger, Evan +7 · 35 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  9. Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
    2024/06/14 by Denison, Carson, MacDiarmid, Monte, Barez, Fazl +11 · 24 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  10. Auditing language models for hidden objectives
    2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 16 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  11. Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
    2023/07/17 by Ansh Radhakrishnan, Radhakrishnan, Ansh, Karina Nguyen +45 · 5 citations
    Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  12. Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
    2024/04/25 by Olli Järviniemi, Järviniemi, Olli, Evan Hubinger +1 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering Research #Topic Modeling
  13. Engineering Monosemanticity in Toy Models
    2022/11/16 by Adam S. Jermyn, Jermyn, Adam S., Nicholas Schiefer +3 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  14. An overview of 11 proposals for building safe advanced AI
    2020/12/04 by Evan Hubinger, Hubinger, Evan · 3 citations
    Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Software Engineering Research
  15. Sabotage Evaluations for Frontier Models
    2024/10/28 by Benton, Joe, Wagner, Misha, Christiansen, Eric +13 · 6 citations
    #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  16. Agentic Misalignment: How LLMs Could Be Insider Threats
    2025/10/05 by Aengus Lynch, Benjamin Wright, Lynch, Aengus +12 · 20 citations
    Business, Management and Accounting · Computer Science · #Securities Regulation and Market Practices #Corporate Insolvency and Governance #Cybercrime and Law Enforcement Studies
  17. Natural Emergent Misalignment from Reward Hacking in Production RL
    2025/11/23 by MacDiarmid, Monte, Wright, Benjamin, Uesato, Jonathan +19 · 8 citations
    Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Software Engineering Research
  18. Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
    2025/11/03 by Sharan Maiya, Maiya, Sharan, Henning Bartsch +5 · 2 voices · 2 citations
    Computer Science · Psychology · #cs.CL #cs.AI #cs.LG
  19. Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
    2025/05/20 by Chiu, Yu Ying, Wang, Zhilin, Maiya, Sharan +4 · 6 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
  20. Conditioning Predictive Models: Risks and Strategies
    2023/02/02 by Evan Hubinger, Hubinger, Evan, Adam S. Jermyn +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling