vix.ing · top · new · best · stats · spec

Wijk, Hjalmar

  1. Measuring AI Ability to Complete Long Software Tasks
    2025/03/18 by Thomas Kwa, Kwa, Thomas, Ben West +48 · 26 voices · 38 citations
    #cs.AI #cs.LG
  2. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
    2024/11/22 by Hjalmar Wijk, Wijk, Hjalmar, Tao Lin +47 · 3 voices · 10 citations
    Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.LG
  3. Evaluating Language-Model Agents on Realistic Autonomous Tasks
    2023/12/18 by Kinniment, Megan, Sato, Lucas Jun Koba, Du, Haoxing +10 · 10 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)