2025/09/30 by Bertie Vidgen, Abby Fennelly, Vidgen, Bertie +37 · 3 voices · 4 citations
Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Economics and business #General Economics (econ.GN) #Human-Computer Interaction (cs.HC) #cs.AI #cs.CL #cs.HC #econ.GN
paper · pdf · doi:10.48550/arxiv.2509.25721
arxiv published 2025/09/30 · arxiv updated 2025/12/16
We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.