Belrose, Nora
- Adversarial Policies Beat Superhuman Go AIs
2022/11/01 by Tony Tong Wang, Tony T. Wang, Wang, Tony T. +21 · 18 voices · 7 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Eliciting Latent Predictions from Transformers with the Tuned Lens
2023/03/14 by Nora Belrose, Igor Ostrovsky, Belrose, Nora +13 · 3 voices · 52 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- LEACE: Perfect linear concept erasure in closed form
2023/06/06 by Nora Belrose, Belrose, Nora, David Schneider-Joseph +9 · 1 voice · 25 citations
Computer Science · Social Sciences · #Machine Learning in Healthcare #Misinformation and Its Impacts #Topic Modeling #cs.CL #cs.CY #cs.LG
- Automatically Interpreting Millions of Features in Large Language Models
2024/10/17 by Gonçalo Paulo, Paulo, Gonçalo, Alex Mallen +6 · 1 voice · 30 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Sparse Autoencoders Trained on the Same Data Learn Different Features
2025/01/28 by Gonçalo Paulo, Nora Belrose, Paulo, Gonçalo +1 · 19 citations
Computer Science · #Machine Learning and Data Classification
- Neural Networks Learn Statistics of Increasing Complexity
2024/02/06 by Belrose, Nora, Pope, Quintin, Quirke, Lucia +2 · 7 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- imitation: Clean Imitation Learning Implementations
2022/11/22 by Gleave, Adam, Taufeeque, Mohammad, Rocamonde, Juan +7 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Does Transformer Interpretability Transfer to RNNs?
2024/04/09 by Gonçalo Paulo, Thomas Marshall, Paulo, Gonçalo +3 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Eliciting Latent Knowledge from Quirky Language Models
2023/12/02 by Alex Mallen, Mallen, Alex, Brumley, Madeline +2 · 4 citations
Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Refusal in LLMs is an Affine Function
2024/11/13 by Thomas Märshall, Thomas Marshall, Marshall, Thomas +4 · 1 voice · 3 citations
Computer Science · #Multi-Agent Systems and Negotiation #cs.CL #cs.LG
- Transcoders Beat Sparse Autoencoders for Interpretability
2025/01/31 by Paulo, Gonçalo, Shabalin, Stepan, Belrose, Nora · 5 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Balancing Label Quantity and Quality for Scalable Elicitation
2024/10/17 by Alex Mallen, Mallen, Alex, Nora Belrose +1 · 1 citation
Pharmacology, Toxicology and Pharmaceutics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Pharmacy and Medical Practices
- Estimating the Probability of Sampling a Trained Neural Network at Random
2025/01/31 by Scherlis, Adam, Belrose, Nora · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Mechanistic Anomaly Detection for "Quirky" Language Models
2025/04/09 by Johnston, David O., Chakraborty, Arkajyoti, Belrose, Nora · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)