vix.ing · top · new · best · stats · spec

Kwa, Thomas

  1. Measuring AI Ability to Complete Long Software Tasks
    2025/03/18 by Thomas Kwa, Ben West, Kwa, Thomas +48 · 26 voices · 39 citations
    #cs.AI #cs.LG
  2. The Singapore Consensus on Global AI Safety Research Priorities
    2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  3. HCAST: Human-Calibrated Autonomy Software Tasks
    2025/03/21 by David B. Rein, Rein, David, Becker, Joel +38 · 3 citations
    Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning
  4. InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
    2024/07/19 by Gupta, Rohan, Arcuschin, Iván, Kwa, Thomas +1 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  5. Compact Proofs of Model Performance via Mechanistic Interpretability
    2024/06/17 by Jason N. Gross, Gross, Jason, Rajashree Agrawal +13 · 1 citation
    Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Model Reduction and Neural Networks