"AI coding tools make developers X% more productive" is the default assumption behind most enterprise AI coding tool purchases. METR (a research organization focused on rigorously measuring AI capabilities) ran a randomized controlled trial in early 2025 that complicates that assumption directly, and the actual result is more interesting, and more useful, than a simple "AI doesn't work" headline would suggest.
What the study actually found
METR's paper, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," assigned real coding tasks to an "AI allowed" or "AI disallowed" condition and had experienced developers record actual completion time under each. The headline result: AI assistance caused tasks to take 19% longer on average, with a confidence interval spanning roughly +2% to +39%, a genuine slowdown, not a null result, on the specific population and task set studied.
Why this doesn't settle the question either way
The nuance matters as much as the headline. For the subset of the original study's developers who participated in a later follow-up, the estimated effect flipped to roughly an 18% speedup, suggesting the tools, the developers' skill in using them, or both, improved meaningfully over the course of 2025. METR's own account of the study notes a real complicating factor for repeating or extending the research: throughout 2025, use of agentic coding tools like Claude Code and Codex increased sharply among open-source developers, which affected both recruitment and retention of study participants, a sign that the underlying tooling landscape was shifting fast enough to complicate direct comparison across the study period.
Why the result doesn't simply generalize to "AI coding tools don't work"
The study specifically measured experienced open-source developers working on tasks in codebases they already knew well, exactly the population and task type where a capable developer's existing deep context is hardest for an AI tool to add much on top of, and where the overhead of reviewing and correcting AI suggestions can plausibly outweigh the time saved generating them. This is consistent with a broader pattern we'd expect: AI coding assistance is most likely to help on less-familiar codebases, more boilerplate-heavy tasks, or for less experienced developers, and least likely to help, or actively slow things down, for exactly the population METR studied.
Where GitHub Copilot's real adoption numbers fit into this picture
None of this has slowed adoption at the market level. GitHub Copilot remains the most widely used tool in the category, with roughly 42% market share among paid AI coding tools and 1.8 million paying subscribers. That gap between a rigorous controlled study finding a productivity cost for one specific population and continued strong market-wide adoption is itself worth sitting with: it suggests either that most real-world usage differs meaningfully from METR's studied population and tasks, that perceived productivity and measured productivity diverge, or both.
What this means for evaluating any productivity claim in this space
Treat any single-number AI coding productivity claim (in either direction) with real skepticism, and evaluate against your own team's specific task mix, developer experience distribution, and codebase familiarity, the same evaluation discipline we recommend in our guide to choosing AI tools and in our step-by-step AI coding assistant setup guide, which explicitly flags this study before recommending a cautious, task-by-task rollout rather than a blanket adoption assumption.
The honest takeaway
METR's finding is a genuinely important data point precisely because it cuts against the industry's dominant narrative, and the responsible response isn't to dismiss it or to treat it as the final word. It's to take seriously that AI coding tool value is real but unevenly distributed, and that "experienced developer working in a familiar codebase" may be exactly the population where the marketing claims are weakest.
