Paulo, Gonçalo
- Automatically Interpreting Millions of Features in Large Language Models
2024/10/17 by Gonçalo Paulo, Paulo, Gonçalo, Alex Mallen +6 · 1 voice · 35 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
2025/05/17 by Guijin Son, Ji-Woo Hong, Son, Guijin +20 · 4 voices · 9 citations
Computer Science · Medicine · Decision Sciences · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Scientific Computing and Data Management
- Sparse Autoencoders Trained on the Same Data Learn Different Features
2025/01/28 by Gonçalo Paulo, Nora Belrose, Paulo, Gonçalo +1 · 20 citations
Computer Science · #Machine Learning and Data Classification
- Does Transformer Interpretability Transfer to RNNs?
2024/04/09 by Gonçalo Paulo, Thomas Marshall, Paulo, Gonçalo +3 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Transcoders Beat Sparse Autoencoders for Interpretability
2025/01/31 by Paulo, Gonçalo, Shabalin, Stepan, Belrose, Nora · 5 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)