METR 2026 Study: What the AI Productivity Data Shows
Key Takeaways (TL;DR)
- 2025: In METR's trial, experienced developers took 19% longer with AI while believing they were 20% faster (METR, 2025).
- 2026: METR's update estimates an 18% cut in task time for 10 returning developers, but the interval runs from -38% to +9%. METR calls the data unreliable.
- The real finding: 30-50% of developers withheld tasks they didn't want to do without AI.
- What to do: Measure your delivery system, not individual speed.
In 2025, experienced open-source developers expected AI to cut their task time by 24%. METR measured the opposite: with AI, their tasks took 19% longer. After the study, they still believed AI had made them 20% faster.
A year later, that study gets quoted in both directions. One camp says "AI slows developers down." The other says "METR now shows an 18% speedup." Neither is what METR says.
If you're deciding what to do about AI coding tools, the difference matters. Below are METR's three releases in order, what each one measured, the three misreadings I see most, and what to measure in your own team instead. For the wider ROI picture, start with DORA's ROI of AI-assisted development findings.
What did METR's 2025 developer productivity study find?
In July 2025, METR published Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, a randomized controlled trial. With AI allowed, tasks took 19% longer, with a confidence interval of +2% to +39% (METR, 2025). The whole interval is above zero, so this was a real slowdown, in one specific setting.
The setup was unusually realistic. 16 experienced developers worked on 246 real issues from their own repositories, each taking about two hours. The repositories averaged more than 22,000 stars and over a million lines of code. Each issue was randomly assigned to "AI allowed" or "AI not allowed." When AI was allowed, developers mostly used Cursor Pro with Claude 3.5 and 3.7 Sonnet.
The number that stuck is the gap between what developers felt and what the clock said. In METR's words:
"This gap between perception and reality is striking: developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%." (METR, 2025)
METR is careful about scope: "We do not claim that our developers or repositories represent a majority or plurality of software development work." These were experts working in code they knew very well, plausibly where an AI assistant has the least to add. So the study shows a slowdown in one setting, plus a perception gap even experts didn't notice.
What changed in METR's February 2026 update?
In February 2026, METR published We are Changing our Developer Productivity Experiment Design. A new experiment, started in August 2025, covered 57 developers, 143 repositories and 800+ tasks. Task time fell an estimated 18% for returning developers and 4% for new ones. Neither result is statistically significant (METR, 2026).
METR reports the change in task time, so a negative number means faster. "-18%" means tasks took 18% less time with AI. The confidence interval for returning developers runs from -38% to +9%. For new developers it runs from -15% to +9%. Both intervals include zero. The data fits a large speedup, no effect, or a small slowdown.
This isn't the 2025 cohort measured again. Only 10 of the original 16 developers came back. The other 47 were recruited from a wider set of open-source projects. METR's post mentions the rise of agentic tools such as Claude Code and Codex during 2025. It doesn't say which tools participants used.
Why does METR call its own 2026 data unreliable?
In its February 2026 update, METR reported: "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI." The tasks developers least wanted to do without AI dropped out, and those are likely where AI helps most. That pushes the measured speedup toward zero.
METR names two more problems with the 2026 data. It cut the pay from $150 to $50 an hour, which it says caused further selection effects. And its time measurements are unreliable for developers who run several AI agents at the same time.
METR's own reading goes further than its numbers. Based on conversations with participants, it believes "it is likely that developers are more sped up from AI tools now" than in its early-2025 estimates. It adds that "the true speedup could be much higher among the developers and tasks which are selected out of the experiment." But METR can't show that with this data. So it's changing the design, moving toward more intensive experiments and observational data.
My read: the selection effect is the finding. When a third to half of your participants hold back tasks rather than do them without AI, "no AI" no longer describes normal work. For these developers, adoption became irreversible faster than a controlled trial could measure it. That has a direct consequence for you. You won't get a clean control group inside your own team either. Nobody will agree to work without AI for eight weeks so you can compare. You have to measure the system in production, before and after the change.
What does METR's 2026 survey add?
In May 2026, METR published a survey of 349 technical workers, run from February to April 2026. The median respondent said AI raised the value of their work 1.4-2x and their speed 3x. They expect 2.5x by March 2027. METR adds its own warning: "survey results are not necessarily grounded in reality."
Respondents recalled 1.3x for March 2025, so the self-reported gains are rising. Two details matter when you quote it. First, only 87 of the 349 respondents were software engineers. The rest were researchers, academics, PhD students, founders and managers. Second, METR's 2025 trial "found that people overestimated AI's effect on their time spent on tasks by 40 percentage points on average."
The most telling subgroup is METR's own staff. They gave the lowest answers of anyone. METR suspects that's because its staff know the earlier findings on perceived versus actual productivity. The people who've seen the perception gap measured are the least impressed. Does that tell you something about self-reported AI gains in your own team?
What most coverage gets wrong about METR
Much of the 2026 coverage of METR repeats one of three misreadings. The one I see most quotes -18% as a result, when its interval runs from -38% to +9%. Each misreading turns a hedged finding into a clean headline. Here's what the primary sources say.
- "METR now shows an 18% speedup." It's a point estimate for 10 returning developers. The interval crosses zero, and METR itself says the speedup could be much higher on the tasks developers held back.
- "The same developers, 12 months later." A new experiment. Only 10 of the original 16 returned, and the tasks, pay and tools on the market had all changed.
- "The survey proves AI doubles productivity." It's self-report, and only a quarter of respondents were software engineers. METR's own trial showed self-estimates overshooting by about 40 percentage points.
I made the first two mistakes myself. My DORA post called the 2026 follow-up a 12-month longitudinal study showing an 18% speedup, until I corrected it on 30 September 2026. The primary sources are short. Read them before you quote a number to your board.
How should engineering leaders measure AI productivity instead?
Stop measuring individual developer speed. METR's 2025 trial showed self-estimates can miss by 40 percentage points, and its 2026 experiment lost tasks from the no-AI condition because developers held them back. Measure the delivery system instead, before and after the rollout. Here are five numbers you can pull from GitHub or GitLab today.
| Metric | What it tells you | Where to get it |
|---|---|---|
| PR cycle time | How long work takes from open to merge | gh pr list --state merged --limit 1000 --search "merged:>=YYYY-MM-DD" --json createdAt,mergedAt |
| Time to first review | Whether review is the new bottleneck | same query, add reviews |
| Median PR size | Whether AI inflates changes past what humans can review | same query, add additions,deletions |
| Revert / change failure rate | Whether speed is costing quality | git log --oneline --grep='^Revert' against merged PRs, or incidents per deploy |
| Deploy frequency | Whether more code turns into more delivery | Your CD tool's deploy log, or gh api repos/{owner}/{repo}/deployments |
On GitLab, the merge requests API (glab api projects/:id/merge_requests) gives you the same fields. Pull eight weeks before the rollout and eight weeks after. Then compare medians, not averages. One giant PR shouldn't decide your AI strategy.
What this looked like in one client team
Here's one of my engagements. It was a corporate startup in the automotive sector, building embedded software. 11 engineers, 2 PMs and 2 QA, working under ASPICE and using Claude Code. Deploys to internal integration happened about once a month.
The team moved to CI/CD and put AI code review, using CodeRabbit, in front of human review as the first gate. Deploys went to roughly twice a week.
I can't give Claude Code or CodeRabbit the credit for that jump. Most of it came from building the automated pipeline. What the AI gate did was let the humans keep up with the higher PR volume.
It wasn't free. The AI reviewer needed fine-tuning. It also set off lots of internal discussions about how much code to review. Those turned out to be discussions about code expectations and style agreements the team had never written down. The AI reviewer forced the implicit rules into the open.
That's why I measure the system: that's where a gain shows up, if it shows up at all. When it doesn't, check the review queue. DORA's data on review bottlenecks covers that failure mode. The AI Jevons Paradox shows the other outcome, where teams ship more but don't save time. And Shift-Up Engineering explains how AI moves engineering work up the stack. To turn these numbers into a business case, use an honest AI coding ROI model.
If your PR queue has grown since you rolled out coding assistants, that's the system I fix as an interim engineering leader. Bring these five numbers to a 30-minute teardown.
What to take from METR's studies
- The 2025 slowdown was real for experts in their own large codebases, and it came with a 40-point perception gap.
- The 2026 speedup is plausible but unproven. Both intervals cross zero, and METR calls the data unreliable.
- Self-reports overshoot. The survey's 1.4-2x value and 3x speed figures are self-reports.
- Clean control groups are getting hard to find. When developers won't do tasks without AI, measure in production.
- The only numbers that count are your own: cycle time, review latency, PR size, failure rate and deploy frequency.
Frequently Asked Questions
Did METR find that AI makes developers slower?
Yes, in early 2025. In METR's randomized trial, 16 experienced open-source developers took 19% longer with AI on their own repositories, with a confidence interval of +2% to +39%. They still believed AI had made them 20% faster. METR cautions that its developers and repositories don't represent most software development.
Does METR's 2026 update show AI makes developers faster?
It points that way, but it doesn't prove it. In February 2026, METR estimated task time fell 18% for returning developers and 4% for new ones. Neither result is statistically significant. METR calls the data unreliable because 30-50% of developers withheld tasks they didn't want to do without AI.
Why did METR change its study design?
Too many developers held back tasks they didn't want to do without AI. In 2026, 30-50% said they skipped submitting some tasks, so the "no AI" condition stopped describing normal work. METR is moving to more intensive experiments and observational data instead of the original randomized design.
How much faster are developers with AI in 2026?
Nobody has a clean measurement. In METR's 2026 survey, 349 technical workers self-reported a median 3x speed gain. METR's controlled data can't confirm it, and its 2025 trial found self-estimates off by about 40 percentage points. Measure your own team's cycle time, review latency and deploy frequency instead.
No affiliation with any tool mentioned in this post.
Sources:
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 2025, retrieved 2026-09-30, https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Becker, Rush, Barnes, Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," arXiv:2507.09089, July 2025, retrieved 2026-09-30, https://arxiv.org/abs/2507.09089
- METR, "We are Changing our Developer Productivity Experiment Design," February 2026, retrieved 2026-09-30, https://metr.org/blog/2026-02-24-uplift-update/
- METR, AI usage survey, May 2026, retrieved 2026-09-30, https://metr.org/blog/2026-05-11-ai-usage-survey/