Kwa, Thomas
- Measuring AI Ability to Complete Long Software Tasks
2025/03/18 by Thomas Kwa, Ben West, Kwa, Thomas +48 · 26 voices · 39 citations
#cs.AI #cs.LG
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- HCAST: Human-Calibrated Autonomy Software Tasks
2025/03/21 by David B. Rein, Rein, David, Becker, Joel +38 · 3 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning
- InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
2024/07/19 by Gupta, Rohan, Arcuschin, Iván, Kwa, Thomas +1 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Compact Proofs of Model Performance via Mechanistic Interpretability
2024/06/17 by Jason N. Gross, Gross, Jason, Rajashree Agrawal +13 · 1 citation
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Model Reduction and Neural Networks