Aidan Ewart
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
2023/09/15 by Hoagy Cunningham, Cunningham, Hoagy, Aidan Ewart +7 · 296 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Adversarial Robustness in Machine Learning
- Eight Methods to Evaluate Robust Unlearning in LLMs
2024/02/26 by Aengus Lynch, Phillip Guo, Lynch, Aengus +7 · 35 citations
Engineering · Health Professions · #Fault Detection and Control Systems #Quality and Safety in Healthcare
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
2024/07/22 by Abhay Sheshadri, Aidan Ewart, Sheshadri, Abhay +19 · 38 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
2025/02/03 by Zora Che, Stephen Casper, Che, Zora +27 · 14 citations
Computer Science · #Security and Verification in Computing #Advanced Malware Detection Techniques #Network Security and Intrusion Detection
- Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
2024/10/16 by Phillip Guo, Guo, Phillip, Syed, Aaquib +6 · 8 citations
Computer Science · #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Neural Networks and Applications