Hoagy Cunningham
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
2023/09/15 by Hoagy Cunningham, Aidan Ewart, Cunningham, Hoagy +7 · 340 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Adversarial Robustness in Machine Learning
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 64 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 23 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG