vix.ing · top · new · best · stats · spec

Euan Ong

  1. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
    2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 44 citations
    Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
  2. Image Hijacks: Adversarial Images can Control Generative Models at Runtime
    2023/09/01 by Luke Bailey, Bailey, Luke, Euan Ong +5 · 25 citations
    Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Auditing language models for hidden objectives
    2025/03/14 by Samuel D. Marks, Samuel Marks, Johannes Treutlein +71 · 1 voice · 17 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  4. Successor Heads: Recurring, Interpretable Attention Heads In The Wild
    2023/12/14 by R. Bruce Gould, Euan Ong, Gould, Rhys +5 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  5. Compact Proofs of Model Performance via Mechanistic Interpretability
    2024/06/17 by Jason N. Gross, Rajashree Agrawal, Gross, Jason +13 · 1 citation
    Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Model Reduction and Neural Networks
  6. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
    2025/12/17 by Adam Karvonen, James Chua, Karvonen, Adam +19 · 1 voice · 1 citation
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
  7. Verbalizable Representations Form a Global Workspace in Language Models
    2026/07/16 by Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13 · 2 voices · 8 citations
    #cs.CL #cs.AI #cs.LG