AI Coding Assistant ROI: How to Calculate It Honestly
Key Takeaways (TL;DR)
- License cost is noise. For 50 engineers, tools at $150 a head cost about $90k a year. A 15% productivity dip for three months costs about $281k in payroll.
- Dip length decides year one. With identical tooling, year-one net value ranges from -$465k to +$972.5k, depending on how long the dip lasts and what gain follows.
- Hours saved aren't dollars saved. Saved time only counts once it turns into shipped work. LinearB found AI-generated PRs wait 4.6x longer for review (LinearB, 2026).
- What to do: Run three scenarios, measure delivery before and after, and spend your effort shortening the dip, not negotiating seat price.
Vendor-commissioned ROI studies for AI coding tools run to several hundred percent. The most-quoted one, "376% ROI", isn't even about Copilot alone. On complex work in existing codebases, the research DORA cites puts the gain at 10% or less.
Then the CFO asks for your number. Almost every calculator online gives the same answer: hours saved per developer times salary, minus the seat price. That model leaves out the costs that actually decide the result.
This post is the model I use instead. It's worked for a 50-engineer team, with three scenarios and every assumption visible, so you can swap in your own numbers. If you want the research behind it first, read my summary of DORA's ROI of AI-assisted development report.
Why do most AI coding ROI numbers overstate the return?
Most AI coding ROI numbers count the seat price as the only cost and self-reported time saved as the benefit. Both inputs are wrong in the same direction. The cost is too low, and the benefit is too high.
Self-reports overshoot. In METR's 2025 randomized trial, experienced open-source developers took 19% longer with AI on their own repositories. Afterwards, they still believed AI had made them 20% faster (METR, 2025). If your benefit line comes from a survey of your own engineers, you're building on the number METR showed was off by about 40 points. I cover what METR's controlled studies found in detail separately.
Gains also depend on the work. Google's DORA team, citing research on tens of thousands of engineers, puts the gain at 35-40% on simple greenfield tasks, but 10% or less on complex brownfield code (Google DORA, 2026). A funded startup a few years in isn't writing greenfield code. Its engineers spend most of the day in an existing codebase.
Then there are the vendor studies. The "376% ROI" figure comes from a Forrester Total Economic Impact study commissioned by GitHub, published in July 2025. It covers GitHub Enterprise Cloud as a platform, with Copilot, Actions and Advanced Security together, for a composite company with 5,000 developers (Forrester, 2025). It's a real study. It just isn't a Copilot ROI number, and it isn't your team.
What does an AI coding assistant actually cost per developer?
List prices for team plans run from $19 to $125 per user per month, plus usage once the included allowance runs out. These are the published prices on 30 September 2026:
| Tool | List price per user/month | Billing model | What's metered |
|---|---|---|---|
| GitHub Copilot Business | $19 | Per seat | 1,900 AI credits per user per month included |
| GitHub Copilot Enterprise | $39 | Per seat | 3,900 AI credits per user per month included |
| Cursor Teams, Standard seat | $40 | Per seat | Included model usage, on-demand usage billed in arrears |
| Cursor Teams, Premium seat | $120 | Per seat | 5x the usage of a Standard seat |
| Claude Team, Standard seat | $20 annual / $25 monthly | Per seat, includes Claude Code | Extra usage can be bought |
| Claude Team, Premium seat | $100 annual / $125 monthly | Per seat, includes Claude Code | Extra usage can be bought |
Sources: GitHub Copilot plans, Cursor pricing, Claude pricing. Prices change often. Check before you quote them.
Teams rarely stop at one seat. A premium agent seat plus usage, or two tools side by side, lands well above the entry price. I use $150 per engineer per month as the planning figure in the model below. For 50 engineers, that's $90,000 a year.
That's the visible cost. The ones that don't show up on an invoice are larger: review time, rework, security review of generated code, governance, and training. DORA's own sample calculator puts licenses at $250 per user per year, plus $80 in extra usage. It then adds $9,600 per user in training and a $3.3M J-curve cost for 500 engineers. In that example, licenses and usage are about 2% of an $8.4M first-year investment (Google DORA, 2026).
How does the J-curve change the calculation?
DORA's 2026 ROI report expects a productivity dip before any gains show up. Its sample calculator uses a 15% drop in engineering capacity for three months. DORA names three causes: the learning curve, the verification tax of reviewing generated code, and pipeline adaptation as testing and approval catch up with the extra volume (Google DORA, 2026).
DORA prices the dip with a simple formula:
J-curve cost = staff size × fully loaded salary × dip % × (dip months / 12)
The duration is the input to watch. Faros AI ran DORA's calculator for a 500-person engineering org. At the default three-month dip, it showed a +39.2% ROI. Stretching the dip to 12 months, with everything else unchanged, produced "a $9.9M swing" and flipped ROI to -36.2% (Faros AI, April 2026).
For a 50-engineer team, month by month, it looks like this:
All three lines fall together for the first two months. What separates them is when the dip ends and how steep the line is afterwards. Neither depends on which tool you bought.
How to calculate AI coding assistant ROI, step by step
Year-one ROI is net value divided by what you put in, and what you put in includes the dip. Here's the formula I use:
- Gain value = payroll × gain % after the dip × (months after the dip / 12)
- Dip cost = payroll × dip % × (dip months / 12)
- Year-one net value = gain value - dip cost - tool cost
- ROI = net value / (tool cost + dip cost)
Here's the worked example: 50 engineers at $150,000 fully loaded, so $7.5M in payroll, and $90,000 a year in tools. The dip is DORA's 15% in every row. Only its length and the gain afterwards change.
| Scenario | Dip | Gain after dip | Gain value | Dip cost | Tools | Year-one net | ROI |
|---|---|---|---|---|---|---|---|
| Conservative | 15% for 6 months | 5% | $187.5k | $562.5k | $90k | -$465k | -71% |
| Realistic | 15% for 3 months | 10% | $562.5k | $281.25k | $90k | +$191.25k | +52% |
| Optimistic | 15% for 2 months | 20% | $1.25M | $187.5k | $90k | +$972.5k | +350% |
Look at the tools column. It's the same in every row. The result swings by $1.4M anyway, because everything that matters sits in the dip and the realized gain.
DORA recommends the same approach: don't present a single number. Build a conservative, a realistic and an optimistic scenario. DORA does it by applying multipliers to value and cost (Google DORA, 2026). I prefer changing dip length and gain directly, because those are the inputs you can check against your own data later.
Now change one input at a time, starting from the realistic scenario:
One extra month of dip costs more than doubling your seat price. That's the whole argument of this post in one chart.
The gain value is not cash. DORA is explicit here. It doesn't assume saved time saves the cost of development. It counts saved time as capacity you reinvest, which saves you future hires. If you model the gain as payroll you'll cut, you're modeling a different decision, one DORA advises against.
What the CEO asked first
In my case, the person asking was the CEO. I was interim Head of Engineering at a corporate startup in the automotive sector: 11 engineers, 2 PMs and 2 QA, building embedded software under ASPICE. The schedule was behind and the budget was fixed. The CEO wanted the team to go faster with AI and asked two questions: is it safe, and does it work?
Neither question was about seat price. "Safe" meant no hidden bugs slipping into automotive code under ASPICE. "Works" meant whether current AI can do this kind of embedded work at all. The team was already writing with Claude Code. The answer to "safe" was more review, not less: the team put an AI reviewer, CodeRabbit, in front of human review as the first gate. The answer to "works" came from delivery data. With the move to CI/CD, deploys to integration went from about once a month to roughly twice a week. Most of that came from the pipeline, not the AI. The seat prices were the cheap part. The system around them decided whether they worked.
Why don't the time savings show up in delivery?
Saved coding time often ends up waiting in the review queue. LinearB analyzed more than 8.1 million pull requests from over 4,800 organizations. AI-generated PRs wait 4.6x longer before a reviewer picks them up, though they're reviewed 2x faster once someone does. Their acceptance rate is 32.7%, against 84.4% for manual PRs (LinearB, 2026).
So a developer who saves two hours writing a change can lose a day waiting for review, and then see it not accepted two times out of three. The keyboard time went down. The time to production didn't. If your ROI model counts keyboard time, it books a gain your customers never see.
The other outcome is more output of the same value. Linear's data from 47,900 workspaces showed teams with coding agents tripling their weekly pull requests while working more, not less (Linear, How Teams Build, June 2026). I call it the AI Jevons Paradox: hours saved turn into more output, not free time. That's only a return if the extra output is work you'd have paid for anyway.
That's why ROI has to be measured on delivery, not keystrokes.
What to measure to prove ROI
Measure the delivery system before and after the rollout, not individual speed. The numbers that belong in an ROI review are the ones finance can connect to shipped work:
- PR cycle time: open to merge. If it grows, the gain is stuck in review.
- Time to first review: the number LinearB's 4.6x lives in.
- Change failure rate: whether speed is being paid for in incidents.
- Deploy frequency: whether more code turns into more releases.
- Share of roadmap shipped: the only one your CFO actually cares about.
Leave lines of code and suggestion acceptance rate out of it. Both go up when AI writes more code, whether or not the code is worth anything.
Pull eight weeks of data before the rollout and eight weeks after, and compare medians. My METR post has the gh and GitLab queries for each metric. Then put your measured numbers back into the formula above, replacing the assumed dip and gain. That turns a forecast into a result you can defend.
How do you shorten the dip?
You shorten the dip by fixing the system the tool plugs into, before and during rollout. DORA's AI Capabilities Model names seven capabilities that amplify AI's effect. They are a clear and communicated AI stance, a healthy data ecosystem, AI-accessible internal data, strong version control practices, working in small batches, a user-centric focus, and a quality internal platform (DORA AI Capabilities Model).
In practice, three moves do most of the work:
- Reserve review capacity before rollout. Decide who reviews AI-generated PRs and how fast, before the volume arrives. Otherwise the 4.6x wait becomes your dip.
- Cap PR size. Small batches are reviewable. A 2,000-line agent PR isn't, and it'll sit.
- Pair the rollout with CI speed work. If your pipeline takes 40 minutes, faster code generation just means a longer queue in front of it.
None of these cost a license. All of them move the numbers in the chart above more than any seat-price negotiation will. For the org-design side, see Shift-Up Engineering on how AI moves engineering work up the stack.
If your ROI case depends on a dip you can't shorten, that's the system I fix as an interim engineering leader. Bring your PR cycle time, review latency, change failure rate and deploy frequency to a 30-minute teardown.
Frequently Asked Questions
How do you calculate the ROI of AI coding assistants?
Net value divided by investment, with the productivity dip counted as a cost. Year-one net value is the payroll value of the gain after the dip, minus the dip's payroll cost, minus tools. For 50 engineers with a 15% three-month dip and a 10% gain afterwards, that's about +$191k, or 52% ROI.
What do AI coding assistants cost per developer in 2026?
Team plans list at $19 to $125 per user per month: GitHub Copilot Business $19, Cursor Teams $40 or $120, and Claude Team premium seats $100-125. Usage beyond the included allowance costs extra. Tooling is still the smallest line: a three-month, 15% dip costs about three times a year of licenses at $150 a head.
How long until AI coding tools pay off?
It depends mostly on how long the dip lasts. In my realistic scenario, a three-month dip followed by a 10% gain, cumulative value turns positive in month 9. With a two-month dip and a 20% gain, it's month 4. With a six-month dip and a 5% gain, it doesn't pay off in year one. DORA's sample calculator shows an eight-month payback.
Is the 376% Copilot ROI figure realistic?
It isn't a Copilot figure. The 376% comes from a GitHub-commissioned Forrester study of GitHub Enterprise Cloud, covering Copilot, Actions and Advanced Security together, for a composite company with 5,000 developers. Treat it as a vendor benchmark for a platform, not an estimate for your team's AI rollout.
No affiliation with any tool mentioned in this post.
Sources:
- Google DORA, "The ROI of AI-assisted Software Development," v. 2026.1, retrieved 2026-09-30, https://services.google.com/fh/files/misc/dora-roi-of-ai-assisted-software-development-2026.pdf
- Faros AI, "DORA ROI of AI: How to Stress-Test It Before Your CFO Sees It," April 2026, retrieved 2026-09-30, https://www.faros.ai/blog/dora-ai-roi-calculator-telemetry-inputs
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 2025, retrieved 2026-09-30, https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Linear, "How Teams Build," June 2026, retrieved 2026-09-30, https://linear.app/data
- LinearB, "2026 Software Engineering Benchmarks Report," retrieved 2026-09-30, https://linearb.io/resources/software-engineering-benchmarks-report
- Forrester, "The Total Economic Impact of GitHub Enterprise Cloud," commissioned by GitHub, July 2025, retrieved 2026-09-30, https://tei.forrester.com/go/github/enterprisecloud/
- DORA, AI Capabilities Model, retrieved 2026-09-30, https://dora.dev/ai/capabilities-model/report/
- GitHub, Copilot plans; Cursor, Teams pricing; Anthropic, Claude pricing, all retrieved 2026-09-30