vix.ing · top · new · best · stats · spec

Paulo, Gonçalo

  1. Automatically Interpreting Millions of Features in Large Language Models
    2024/10/17 by Gonçalo Paulo, Paulo, Gonçalo, Alex Mallen +6 · 1 voice · 35 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  2. When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
    2025/05/17 by Guijin Son, Ji-Woo Hong, Son, Guijin +20 · 4 voices · 9 citations
    Computer Science · Medicine · Decision Sciences · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Scientific Computing and Data Management
  3. Sparse Autoencoders Trained on the Same Data Learn Different Features
    2025/01/28 by Gonçalo Paulo, Nora Belrose, Paulo, Gonçalo +1 · 20 citations
    Computer Science · #Machine Learning and Data Classification
  4. Does Transformer Interpretability Transfer to RNNs?
    2024/04/09 by Gonçalo Paulo, Thomas Marshall, Paulo, Gonçalo +3 · 1 voice · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  5. Transcoders Beat Sparse Autoencoders for Interpretability
    2025/01/31 by Paulo, Gonçalo, Shabalin, Stepan, Belrose, Nora · 5 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)