Euan Ong
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 44 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime
2023/09/01 by Luke Bailey, Bailey, Luke, Euan Ong +5 · 25 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Johannes Treutlein +71 · 1 voice · 17 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Successor Heads: Recurring, Interpretable Attention Heads In The Wild
2023/12/14 by R. Bruce Gould, Euan Ong, Gould, Rhys +5 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Compact Proofs of Model Performance via Mechanistic Interpretability
2024/06/17 by Jason N. Gross, Rajashree Agrawal, Gross, Jason +13 · 1 citation
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Model Reduction and Neural Networks
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
2025/12/17 by Adam Karvonen, James Chua, Karvonen, Adam +19 · 1 voice · 1 citation
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- Verbalizable Representations Form a Global Workspace in Language Models
2026/07/16 by Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13 · 2 voices · 8 citations
#cs.CL #cs.AI #cs.LG