Emmanuel Ameisen
- Auditing language models for hidden objectives
2025/03/14 by Samuel Marks, Samuel D. Marks, Johannes Treutlein +71 · 1 voice · 16 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- When Models Manipulate Manifolds: The Geometry of a Counting Task
2026/01/08 by Wes Gurnee, Emmanuel Ameisen, Isaac Kauvar +4 · 4 voices · 2 citations
#cs.LG
- Mechanisms of Introspective Awareness
2026/03/22 by Uzay Macar, Li Yang, Atticus Wang +3 · 4 voices
#cs.LG
- Verbalizable Representations Form a Global Workspace in Language Models
2026/07/16 by Wes Gurnee, Nicholas Sofroniew, Adam Pearce +13 · 2 voices · 8 citations
#cs.CL #cs.AI #cs.LG