vix.ing · top · new · best · stats · spec

Evan Hubinger

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  3. Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
    2023/08/07 by Roger Grosse, Juhan Bae, Grosse, Roger +31 · 3 voices · 49 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
  4. Alignment faking in large language models
    2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 63 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  5. Measuring Faithfulness in Chain-of-Thought Reasoning
    2023/07/17 by Tamera Lanham, Lanham, Tamera, Anna Chen +59 · 4 voices · 82 citations
    Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
  6. Steering Llama 2 via Contrastive Activation Addition
    2023/12/09 by Nick Gabrieli, Panickssery, Nina, Gabrieli, Nick +8 · 158 citations
    Computer Science · #Interactive and Immersive Displays
  7. Discovering Language Model Behaviors with Model-Written Evaluations
    2022/12/19 by Ethan Perez, Sam Ringer, Perez, Ethan +123 · 120 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
  8. Risks from Learned Optimization in Advanced Machine Learning Systems
    2019/06/05 by Evan Hubinger, Chris van Merwijk, Hubinger, Evan +7 · 34 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  9. Auditing language models for hidden objectives
    2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 15 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  10. Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
    2023/07/17 by Ansh Radhakrishnan, Radhakrishnan, Ansh, Karina Nguyen +45 · 5 citations
    Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  11. Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
    2024/04/25 by Olli Järviniemi, Järviniemi, Olli, Evan Hubinger +1 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering Research #Topic Modeling
  12. Engineering Monosemanticity in Toy Models
    2022/11/16 by Adam S. Jermyn, Nicholas Schiefer, Jermyn, Adam S. +3 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  13. An overview of 11 proposals for building safe advanced AI
    2020/12/04 by Evan Hubinger, Hubinger, Evan · 3 citations
    Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Software Engineering Research
  14. Agentic Misalignment: How LLMs Could Be Insider Threats
    2025/10/05 by Aengus Lynch, Lynch, Aengus, Benjamin Wright +12 · 20 citations
    Business, Management and Accounting · Computer Science · #Securities Regulation and Market Practices #Corporate Insolvency and Governance #Cybercrime and Law Enforcement Studies
  15. Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
    2025/11/03 by Sharan Maiya, Maiya, Sharan, Henning Bartsch +5 · 2 voices · 2 citations
    Computer Science · Psychology · #cs.CL #cs.AI #cs.LG
  16. Conditioning Predictive Models: Risks and Strategies
    2023/02/02 by Evan Hubinger, Hubinger, Evan, Adam S. Jermyn +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling