McGrath, Thomas
- Tracr: Compiled Transformers as a Laboratory for Interpretability
2023/01/12 by David Lindner, János Kramár, Lindner, David +9 · 1 voice · 5 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #cs.AI #cs.LG #stat.ML
- The Hydra Effect: Emergent Self-repair in Language Model Computations
2023/07/28 by McGrath, Thomas, Rahtz, Matthew, Kramar, Janos +2 · 12 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Copy Suppression: Comprehensively Understanding an Attention Head
2023/10/06 by McDougall, Callum, Conmy, Arthur, Rushing, Cody +2 · 6 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)