vix.ing · top · new · best · stats · spec

Garriga-Alonso, Adrià

  1. Towards Automated Circuit Discovery for Mechanistic Interpretability
    2023/04/28 by Arthur Conmy, Augustine N. Mavor-Parker, Conmy, Arthur +7 · 79 citations
    Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG)
  2. Bayesian Neural Network Priors Revisited
    2021/02/12 by Fortuin, Vincent, Garriga-Alonso, Adrià, Ober, Sebastian W. +5 · 9 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  3. Deep Convolutional Networks as shallow Gaussian Processes
    2018/08/16 by Adrià Garriga-Alonso, Carl Edward Rasmussen, Garriga-Alonso, Adrià +3 · 11 citations
    Computer Science · #Advanced Neural Network Applications #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  4. Exact Langevin Dynamics with Stochastic Gradients
    2021/02/02 by Garriga-Alonso, Adrià, Fortuin, Vincent · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  5. Among Us: A Sandbox for Measuring and Detecting Agentic Deception
    2025/04/05 by Satvik Golechha, Adrià Garriga-Alonso, Golechha, Satvik +1 · 6 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Deception detection and forensic psychology #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  6. Hypothesis Testing the Circuit Hypothesis in LLMs
    2024/10/16 by Claudia Shi, Shi, Claudia, Nicolas Beltran-Velez +13 · 1 voice · 4 citations
    Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.AI #cs.LG #stat.ML
  7. Planning in a recurrent neural network that plays Sokoban
    2024/07/22 by Mohammad Taufeeque, Taufeeque, Mohammad, Philip Quirke +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  8. Understanding Variational Inference in Function-Space
    2020/11/18 by Burt, David R., Ober, Sebastian W., Garriga-Alonso, Adrià +1 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  9. Data augmentation in Bayesian neural networks and the cold posterior effect
    2021/06/10 by Nabarro, Seth, Ganev, Stoil, Garriga-Alonso, Adrià +3 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  10. Interpreting Emergent Planning in Model-Free Reinforcement Learning
    2025/04/02 by Bush, Thomas, Chung, Stephen, Anwar, Usman +2 · 3 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  11. Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
    2025/05/16 by D.I. Chanin, Chanin, David, Tomáš Dulka +3 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques
  12. InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
    2024/07/19 by Gupta, Rohan, Arcuschin, Iván, Kwa, Thomas +1 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)