Ben West
- Measuring AI Ability to Complete Long Software Tasks
2025/03/18 by Thomas Kwa, Ben West, Kwa, Thomas +48 · 26 voices · 39 citations
#cs.AI #cs.LG
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
2024/11/22 by Hjalmar Wijk, Wijk, Hjalmar, Tao Lin +47 · 3 voices · 10 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.LG