Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
2025/07/12 by Joel Becker, Becker, Joel, Nate Rush +7 · 93 voices · 55 citations
Business, Management and Accounting · Decision Sciences · #A priori and a posteriori #Big Data and Business Intelligence #Function (biology) #Human multitasking #Productivity #Quality (philosophy) #Robustness (evolution) #Scientific Computing and Data Management #Slowdown #Software #Task (project management)
paper · pdf · doi:10.48550/arxiv.2507.09089
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/07/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI tools at the February-June 2025 frontier affect the productivity of experienced open-source developers. 16 developers with moderate AI experience complete 246 tasks in mature projects on which they have an average of 5 years of prior experience. Each task is randomly assigned to allow or disallow usage of early 2025 AI tools. When AI tools are allowed, developers primarily use Cursor Pro, a popular code editor, and Claude 3.5/3.7 Sonnet. Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowed developers down. This slowdown also contradicts predictions from experts in economics (39% shorter) and ML (38% shorter). To understand this result, we collect and evaluate evidence for 20 properties of our setting that a priori could contribute to the observed slowdown effect--for example, the size and quality standards of projects, or prior developer experience with AI tooling. Although the influence of experimental artifacts cannot be entirely ruled out, the robustness of the slowdown effect across our analyses suggests it is unlikely to primarily be a function of our experimental design.
Citations
Cited by
- Enhancing SLMs for Sustainable Code Optimization in Radio-Astronomy
- Pomona: Continuous Code Quality Improvement via Small, Agentic Pull Requests at Bloomberg
- Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
- Position: Natural Language Should Not Fully Replace Formal Languages
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- How AI Impacts Skill Formation
- Speed at the Cost of Quality
- Ten Simple Rules for AI-Assisted Coding in Science
- "Maybe We Need Some More Examples:" Individual and Team Drivers of Developer GenAI Tool Use
- To Ban or not to Ban? How Open Source Projects Govern GenAI Contributions
- Stand-Alone Complex or Vibercrime? Exploring the adoption and innovation of GenAI tools, coding assistants, and agents within cybercrime ecosystems
- Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- An Empirical Study of Generative AI Adoption in Software Engineering
- The Hitchhiker's Guide to Monoculture
- How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- A survey of generative AI adoption and perceived productivity among scientists who program
- AI Researchers' Views on Automating AI R&D and Intelligence Explosions
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
- Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
- Validity Is What You Need
- AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertise in Hybrid Intelligence Systems
- User Misconceptions of LLM-Based Conversational Programming Assistants
- AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
- The efficiency-gain illusion: People underestimate the rate of AI use and overestimate its benefits on simple tasks
- The Fast and Spurious: Developer Productivity with GenAI
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- Product Manager Practices for Delegating Work to Generative AI: "Accountability must not be delegated to non-human actors"
- Vibe Coding: Toward an AI-Native Paradigm for Semantic and Intent-Driven Programming
- Modeling Developer Burnout with GenAI Adoption
- Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
- Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications in AI-assisted Data Analysis
- Intuition to Evidence: Measuring AI's True Impact on Developer Productivity
- Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
- Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
- Vibe Coding for UX Design: Understanding UX Professionals' Perceptions of AI-Assisted Design and Development
- Revolution or Hype? Seeking the Limits of Large Models in Hardware Design
- Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop
- Collaborating with GenAI: Incentives and Replacements
- On the Future of Software Reuse in the Era of AI Native Software Engineering
- SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- Automation, AI, and the Intergenerational Transmission of Knowledge
- Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
- The Scaling Paradox in Human-AI Collaboration
- A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents
- Flaws in the LLM Automation Narrative
- Life After Benchmark Saturation: A Case Study of CORE-Bench
- The Semi-Executable Stack: Agentic Software Engineering and the Expanding Scope of SE
- RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies
- Measuring AI R&D Automation
- Agentic Much? Adoption of Coding Agents on GitHub
- Dynamic Memory Management on GPUs with SYCL
- From Horizontal Layering to Vertical Integration: A Comparative Study of the AI-Driven Software Development Paradigm
- Changes in Coding Behavior and Performance Since the Introduction of LLMs
- Making AI Visible, Not Vanished: How AI Policies Reshape Developer Experience on GitHub
Discussions
- a.i. in fact makes developers less efficient (while perceiving themselves as more efficient) [bsky, 39 points, 3 comments]
- The most recent survey on the use of AI in programming showed a decrease in production speed by 19%, though the programmers thought it increased their speed by 20%. The success of AI is to make people [bsky, 37 points, 2 comments]
- saw a post suggesting that "ai does 70% of the work for you" but all i can think about is the study that found "using ai makes you 19% slower" arxiv.org/abs/2507.09089 so if you handwave the math: ai [bsky, 36 points, 1 comments]
- This study shows that vibe coding makes people code *more slowly* even though they feel like they're coding more quickly arxiv.org/abs/2507.09089 [bsky, 26 points, 3 comments]
- I keep thinking about this study showing programmers who used LLM assistance *believed* that doing so reduced the time to completion by about 20%. However, on average they took 19% longer than the pro [bsky, 17 points, 1 comments]
- mfw the evidence is clear arxiv.org/abs/2507.09089 [bsky, 16 points, 0 comments]
- arxiv.org/abs/2507.09089 [bsky, 14 points, 1 comments]
- New Cornell study: Experienced developers were 19% slower when using AI tools like Claude 3.5 and Cursor Pro—despite expecting 20% time savings. The reason? Low reliability of current AI assistants in [bsky, 12 points, 2 comments]
- Not AI users vs general population, but there was a study on dev productivity in/re how productive they thought they were being vs how productive they were actually being. [bsky, 12 points, 0 comments]
- "After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowed de [bsky, 11 points, 1 comments]
- Yeah, but does it really? My understanding is that research has demonstrated it actually takes more time to use AI tools and fix its mistakes than to just have a programmer do it. arxiv.org/abs/2507.0 [bsky, 9 points, 3 comments]
- AI slows down experienced developers by 19%, this study shows, while they estimated to be much faster with it. They had 16 developers complete 246 tasks in mature open-source projects on which they ha [bsky, 8 points, 0 comments]
- arxiv.org/abs/2507.09089 [bsky, 8 points, 0 comments]
- Off slightly, it was 20% and 19%. arxiv.org/abs/2507.09089 [bsky, 7 points, 0 comments]
- Counterpoint to this anecdote: arxiv.org/abs/2507.09089 “Allowing AI actually increases completion time by 19%--AI tooling slowed developers down.” [bsky, 7 points, 1 comments]
- Yes, the paper studying the effects said the devs involved also estimated that they were more efficient - they expected to be 25% more efficient and thought they were 20% more efficient after doing it [bsky, 7 points, 2 comments]
- Be careful you're not just fooled by the wow-effect, and accidentally making yourself a noticeably worse programmer. There's research even experienced professionals who think they are more effective u [bsky, 5 points, 2 comments]
- Are developers slower using AI? Pro Tip: Read the discussion. p10f Yes, slower for experienced developers and a well known code base. Greenfield or less experience are a different story. And things wi [bsky, 5 points, 1 comments]
- There is evidence of people feeling like it improved efficiency, but in reality the effect was the opposite. arxiv.org/abs/2507.09089 [bsky, 5 points, 0 comments]
- lol, lmao even. Forecasts: developers and ML experts forecast tasks would take significantly less time using LLM tools, *even after* personally participating in the study. Actual results: it's actuall [bsky, 5 points, 0 comments]
- huh "After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowe [bsky, 4 points, 1 comments]
- Measuring the Impact of Early-2025 AI on Experienced Developer Productivity [hn, 4 points, 2 comments]
- arxiv.org/abs/2507.09089 [bsky, 3 points, 0 comments]
- Fantastic paper from METR on the true impact of gen AI on development time. This study shows that experienced developers are less productive when using generative models than not. Specifically, on com [bsky, 3 points, 1 comments]
- Somewhat skeptical of my ability to guess that accurately because of this study showing that devs can overestimate, but maybe "makes coding feel more efficient even if it isn't" actually has a tiny bi [bsky, 3 points, 2 comments]
- Have you seen the preprint on the topic already, where the programmers thought AI made them faster, but it actually made them slower? arxiv.org/abs/2507.09089 [bsky, 3 points, 1 comments]
- The really fun part of this research is that, for at least one study, the subject engineers anticipated a 24% increase in productivity, but the LLMs actually *slowed task completion by 19%*. arxiv.org [bsky, 3 points, 0 comments]
- "Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surpri [bsky, 3 points, 1 comments]
- I'll show you mine if you show me yours: arxiv.org/abs/2507.09089 [bsky, 2 points, 2 comments]
- This one? arxiv.org/abs/2507.09089 [bsky, 2 points, 0 comments]
- all things considered, I think it's important to make sure we keep this bit (from the actual paper) in mind arxiv.org/pdf/2507.090... [bsky, 2 points, 0 comments]
- Maybe as this: arxiv.org/abs/2507.09089 says "Furthermore, we show that both experts and developers drastically overestimate the usefulness of AI on developer productivity, even after they have spent [bsky, 2 points, 1 comments]
- There's a study showing interesting effects in experienced coders, using ai to help them. arxiv.org/abs/2507.09089 Tl;dr - a majority think it *will* make them more productive; a smaller majority thin [bsky, 2 points, 1 comments]
- Measuring the Impact of Early-2025 AI on Experienced Developer Productivity [hn, 2 points, 0 comments]
- AI singularity is just marketing nonsense. AI doesn't even really increase productivity. It just creates the *illusion* of productivity. arxiv.org/abs/2507.09089 [bsky, 2 points, 1 comments]
- arxiv.org/abs/2507.09089 In real life people do actually spend more time fixing the Generative AI slop than if they did the task themselves More time to "lay out in detail the goals and methods" only [bsky, 2 points, 2 comments]
- independent research says yes. Ai increases completion time = lower productivity. And the gap between industry hype (“expert”) predicted increases in productivity and actual decrease in productivity i [bsky, 2 points, 0 comments]
- Ah yep, arxiv.org/abs/2507.09089 was the paper I was thinking of, thanks @anarcish.bsky.social for linking. Obviously I've come to quite a different conclusion about it than you did but yeah. [bsky, 2 points, 1 comments]
- arxiv.org/abs/2507.09089 This study shows it makes you about...19%.....slower than just doing it without AI, LOL. [bsky, 2 points, 0 comments]
- TL/DR Experienced developers predicted AI would decrease time to complete tasks by 24%. Afterwards, those same developers estimated AI had decreased time to complete tasks by 20%. Actual measurements [bsky, 2 points, 0 comments]
- This study seems to show that using genAI actually slows down experienced developers. arxiv.org/abs/2507.09089 [bsky, 2 points, 1 comments]
- My best experience there was the paper which measured coding and while devs thought they were faster (estimated 24% before start, 20% after doing), they were actually 19% *slower*. Perception did not [bsky, 1 points, 0 comments]
- Study on LLM for programming: "Developers forecast AI will reduce completion time by 24%. We find that allowing AI increases completion time by 19%--AI tooling slowed developers down." arxiv.org/abs/2 [bsky, 1 points, 1 comments]
- arxiv.org/abs/2507.09089 There are a handful of papers on it [bsky, 1 points, 1 comments]
- Também tem o detalhe de que, para algumas tarefas, até há a _percepção_ de ganho de eficiência, mas ela não necessariamente corresponde à realidade. Segundo uma pesquisa, programadores se sentem 20% m [bsky, 1 points, 0 comments]
- Can you qualify "largely"? I'm only aware of one study (METR) which was pretty limited in its findings/methodology (mostly discussed in appendix B) and there has been a lot of development in the meant [bsky, 1 points, 1 comments]
- If they want a dense read to ignore and throw in the garbage arxiv.org/pdf/2507.09089 [bsky, 1 points, 0 comments]
- arxiv.org/abs/2507.09089 [bsky, 1 points, 0 comments]
- I passed this along to my scrummaster at work, and he barely reacted... But Cornell did a study that shows programmers who use AI think they're output has improved by 20%, but it actually decreased by [bsky, 1 points, 1 comments]
- Me too arxiv.org/pdf/2507.09089 [bsky, 1 points, 0 comments]
- Like "productive" is something we can feel, and validity doesn't apply to feelings, and also something we can measurably be. Something like this: arxiv.org/abs/2507.09089 [bsky, 1 points, 1 comments]
- Thats kind of part of the issue. What they did for their experiment isn't straightforward at all, and the fact is *looks* that way is a problem. So, the thing you linked above isn't the study either - [bsky, 1 points, 1 comments]
- There we go! 💥 Science has it what I was saying for years: Turns out real-live applications are more complicated than a toy project a first semester can solve 😅 arxiv.org/abs/2507.09089 [bsky, 1 points, 0 comments]
- arxiv.org/abs/2507.09089 [bsky, 1 points, 0 comments]
- But not the majority. For most, it's a straight up loss. I linked one study here. There are several more. I exclude anthropic here, as they're an outlier, and make misleading statements like: "Potenti [bsky, 1 points, 1 comments]
- Compare the title choice here as a more responsible way of communicating such findings: arxiv.org/abs/2507.09089 [bsky, 1 points, 1 comments]
- and notice that the error bars overlap with zero and "However, we are underpowered to draw strong conclusions from this analysis." it's a great study! arxiv.org/pdf/2507.09089 [bsky, 1 points, 1 comments]
- METR (2025) — cautionary tale opener. arxiv.org/abs/2507.09089 An example of a study where they actually collected data on outcomes that compete with task completion to happen first, but then ignored [bsky, 1 points, 1 comments]
- Measuring Impact of Early-2025 AI on Experienced Open-Source Dev Productivity [hn, 1 points, 0 comments]
- Measuring the Impact of AI on Experienced Open-Source Developer Productivity [hn, 1 points, 1 comments]
- YOU can , but could a beginner? Also, you might be an outlier. That study from last summer found people thought they were faster but were actually about 20% slower. [bsky, 1 points, 1 comments]
- Huh... what if experienced programmers were actually getting worse at their job when assisted by LLMs? The possibility that AI-assisted programming just "feels" more efficient is genuine. You might be [bsky, 1 points, 0 comments]
- Imma just leave this here for context. [bsky, 1 points, 0 comments]
- 今日の #AWSSummit で出てきた論文、ChatGPTにまとめてもらった。 ・「熟練者 × 慣れた巨大コードベース」では、生成AI利用で所要時間はむしろ伸びた。 ・作業者自身も「生成AIで20%は短縮した」と錯覚していた。 …あたりが興味深い。 arxiv.org/pdf/2507.09089 [bsky, 0 points, 1 comments]
- arxiv.org/pdf/2507.090... [bsky, 0 points, 0 comments]
- How odd then that a recent random-controlled study *in your field* showed perceived time savings by LLM users were *actually* time sinks. Despite believing they were performing tasks ~20% faster, LLM [bsky, 0 points, 1 comments]
- Here’s that research I was referencing, shows that experienced coders worked up to 19% slower on task they understood. I think the biggest thing here, is that before the test they were confident they [bsky, 0 points, 1 comments]
- arxiv.org/abs/2507.09089 [bsky, 0 points, 0 comments]
- Prompted by my colleagues sharing a news article, I tracked down the paper that the article was based on. Fascinating results. It seem on the surface that 'vibe coding' is a waste of resources. arxiv. [bsky, 0 points, 0 comments]
- Are they though? If you actually know the language strongly, they just slow you down, by nearly 20%. There’s been multiple studies backing up that conclusion: arxiv.org/abs/2507.09089 [bsky, 0 points, 1 comments]
- Read the paper I referenced before speculating on what is going on. They studied experienced engineers working on real tasks in mature codebases. arxiv.org/abs/2507.09089 [bsky, 0 points, 1 comments]
- The full paper is available here arxiv.org/abs/2507.09089 [bsky, 0 points, 0 comments]
- Related, people thought they were faster when they measurably were not: "After completing ... developers estimate ... AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually [bsky, 0 points, 0 comments]
- Typical AI shill, just making shit up. Meanwhile, over here in reality, it turns out that the "transformation" is all in your head, and it actually makes developers LESS productive by the same amount [bsky, 0 points, 0 comments]
- On the other hand, are 16 developers enough population size to claim that AI slows down development? arxiv.org/abs/2507.090... [bsky, 0 points, 1 comments]
- On the other hand, are 16 developers enough population size to claim that AI slows down development? https://arxiv.org/abs/2507.09089v1 [bsky, 0 points, 0 comments]
- #ContraAI 🤖 arxiv.org/pdf/2507.09089 A recent study has been “Measuring the Impact of Early-2025 #AI on Experienced Open-Source Developer Productivity.” The result: “AI tooling slowed developers down [bsky, 0 points, 1 comments]
- arxiv.org/abs/2507.09089 futurism.com/artificial-i... www.economist.com/finance-and-... futurism.com/artificial-i... hbr.org/2026/02/ai-d... (for this one I dont see weekend work jumping by 50% as pro [bsky, 0 points, 1 comments]
- On GenAI as a "Productivity Amplifier": "Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowed developers down." arxiv.org/abs/2507.09089 [bsky, 0 points, 0 comments]
- realmente aumenta? que pesquisa diz isso? só se for alguma pesquisa feita em conjunto com empresa de ia que quer vender a ideia de produtividade. arxiv.org/abs/2507.09089 hbr.org/2025/09/ai-g... [bsky, 0 points, 1 comments]
- Full paper: arxiv.org/abs/2507.09089 [bsky, 0 points, 0 comments]
- @larianstudios.com I want to point this research out for your CEO: arxiv.org/abs/2507.09089 The chances are, that AI is gonna make things slower. Just do yourselves a favour and abandon it, fellars. [bsky, 0 points, 0 comments]
- Surprising slowdown for senior devs due to AI? 📉 A METR RCT tested top AI tools (Cursor Pro, Claude 3.5/3.7) on experienced open-source maintainers, revealing a 19% slowdown in real-world tasks. The [bsky, 0 points, 0 comments]
- Alerted by Albert Edwards to this study of AI's impact on software development. Those taking part thought AI would reduce completion time by 24%; actually, it INCREASED by 19%. Is all that AI spending [bsky, 0 points, 0 comments]
- ok i sent https://arxiv.org/pdf/2507.09089 to the parents lets see how they react haha react [bsky, 0 points, 0 comments]
- Parece que não muito, dependendo da empresa de tecnologia pode até ter piorado a produtividade segundo alguns estudos: arxiv.org/abs/2507.090... www.terra.com.br/noticias/edu... [bsky, 0 points, 1 comments]
- Paper for interest: "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", arxiv.org/abs/2507.09089 TL;DR: the impact was negative [bsky, 0 points, 0 comments]
- arxiv.org/abs/2507.090... [bsky, 0 points, 0 comments]
- The science is catching up... slowly. "After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases comple [bsky, 0 points, 0 comments]
- But Mr. UltimApe, we have studies showing LLMs slow down developers!!! Yeah. and I've read it. Developers are spending time waiting for Cursor / Claude to output code instead of making it a background [bsky, 0 points, 1 comments]
- Source for the 20% productivity loss: arxiv.org/abs/2507.09089 [bsky, 0 points, 0 comments]
- So this is really interesting for a few reasons... arxiv.org/pdf/2507.09089 but mostly because, even as an AI skeptic, there's a bunch of results I found surprising. This study evaluated estimated vs [bsky, 0 points, 1 comments]
- Also, where did you see that only 1/16 had meaningful experience. The cohort had moderate XP it appears: arxiv.org/abs/2507.09089 [bsky, 0 points, 1 comments]
Related