Measuring AI Ability to Complete Long Software Tasks
2025/03/18 by Thomas Kwa, Ben West, Kwa, Thomas +48 · 26 voices · 38 citations
#cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2503.14499
Abstract
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.
Cited by
Discussions
- Exactly. If the decision space increases (larger context), it makes errors more likely. Just read a paper, that reflects this as well: arxiv.org/abs/2503.14499 The "survival rate" increase seems to be [bsky, 12 points, 0 comments]
- Anyways, let's start with the main cite most people push on: arxiv.org/pdf/2503.14499 This is, "Measuring AI Ability to Complete Long Task". The general standard is "how long and complex a task can th [bsky, 10 points, 1 comments]
- Hmmm. What other "categories of failure" does GPT-4 have? I'd iike to see a comprehensive list arxiv.org/pdf/2503.14499 [bsky, 9 points, 2 comments]
- The paper: arxiv.org/pdf/2503.14499 [bsky, 7 points, 0 comments]
- Measuring AI Ability to Complete Long Tasks [hn, 5 points, 0 comments]
- Measuring AI Ability to Complete Long Tasks [hn, 4 points, 0 comments]
- Interesting plot showing how much time a task completed by AI requires for a human professional: about a minute for gpt-2, but around 1 hour for the latest version of Claude.
arxiv.org/abs/2503.14499 [bsky, 4 points, 0 comments]
- Measuring AI Ability to Complete Long Tasks [hn, 3 points, 0 comments]
- For the details, read “Measuring AI Ability to Complete Long Tasks,” now on available on arXiv: arxiv.org/abs/2503.14499 [bsky, 3 points, 1 comments]
- The paper itself does a good job highlighting the limitation. But notice the difference in the plot from the paper vs the plots that are commonly shared. The paper is here: arxiv.org/pdf/2503.14499 [bsky, 2 points, 0 comments]
- lol ty. i legit did try to find a clear definition of "to succeed" re task completion with these different AI models. basically it seems like anthropic is saying 1) claude's own metrics say it's doing [bsky, 2 points, 1 comments]
- Measuring AI Ability to Complete Long Tasks [hn, 2 points, 0 comments]
- ok now you’ve forced me to read their paper. it seems like there’s an actual method here, and it indeed is based on how long a human would take arxiv.org/abs/2503.14499 [bsky, 1 points, 1 comments]
- This paper is important.
arxiv.org/abs/2503.14499
"frontier AI time horizon has been doubling approximately every seven months since 2019... may have accelerated in 2024.... within 5 years, AI syste [bsky, 1 points, 0 comments]
- no like the data is here man. unless you’re a creative, if you’re in the tech sector, get ready to pay for a new degree next year [bsky, 1 points, 2 comments]
- Measuring AI Ability to Complete Long Tasks [hn, 1 points, 0 comments]
- PS: if you're not sure whether AI will ever be capable: its ability to take on more complex and effortful tasks has been doubling every seven months since 2019. See this paper: arxiv.org/abs/2503.1449 [bsky, 0 points, 0 comments]
- 2/ “In a recent METR’s report, the length of coding tasks AIs can handle, their “time horizon”, doubled every 7 months from 2019 - 2024 and every 4 months from 2024-onward... arxiv.org/pdf/2503.14499 [bsky, 0 points, 1 comments]
- “If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take hu [bsky, 0 points, 0 comments]
- [Weekend Read] Measuring AI Ability to Complete Long Tasks - arxiv.org/pdf/2503.14499 The ability of models to perform longer and longer tasks roughly double every 7 months. That's encouraging however [bsky, 0 points, 0 comments]
- The oldest models tested, OpenAI's GPT-2 & 3, can't handle any tasks >1min. The newest, Anthropic's Claude Sonnet 3.7, was 50% successful at tasks that took expert programmers 59 min on avg.
Top mod [bsky, 0 points, 1 comments]
- you can read their methodology here arxiv.org/pdf/2503.14499 [bsky, 0 points, 0 comments]
- METR has released an in-depth study that measures the ability for LLM agents to solve long tasks, and there's quite a lot of interesting insights: Paper: arxiv.org/abs/2503.14499 Post: metr.org/blog/2 [bsky, 0 points, 0 comments]
- Guess what? AI capabilities are following Moore's Law.
arxiv.org/pdf/2503.14499 [bsky, 0 points, 0 comments]
- Measuring AI Ability to Complete Long Tasks [Kwa+, 2025] "50%-task-completion time horizon" measures human experts' time to complete tasks that AI models can complete with a 50% success rate. The time [bsky, 0 points, 0 comments]
- Researchers made a graph with model release date on the x-axis and "task time (for humans)" on the y-axis, and, making the y-axis logarithmic, it forms a straight line (more or less), which means it's [bsky, 0 points, 0 comments]
Related