Skip to main content

Back to Insights Educational article

Measuring AI ROI: How to Get a Number You Can Defend (2026)

If your AI ROI number does not move when you turn the feature off in an A/B test, you do not have ROI. You have a hope dressed up in a slide.

Last updated: May 24, 2026 ROI

~12 min read Reading time 7 Sections Data guide Data-rich article

If your AI ROI number does not move when you turn the feature off in an A/B test, you do not have ROI. You have a hope dressed up in a slide. Before you model anything, anchor on a realistic cost baseline using the AI implementation cost guide.

INTERACTIVE – ROI CALCULATOR

Estimate Year-1 ROI for your AI build

$80K $20K $400K 30% 0% 80% $6K $1K $50K $35K $0K $150K

Total Y1 investment

$176K

Annual value

$420K

Year-1 ROI

+139%

Payback

6 mo

Directional only. A real number needs a written attribution method per layer and a confidence range, not a single point estimate.

01 Reality check

Why most AI ROI dashboards survive only because nobody A/B tests them

The headline number on a CFO deck is almost never measured. It is calculated. Someone counts tickets the bot touched, multiplies by a fully loaded agent cost, subtracts a license fee and writes a ROI percentage that nobody reproduces. The MIT NANDA GenAI Divide report from mid 2025 found that roughly 95% of enterprise generative AI pilots showed no measurable P&L impact at the time of review. The five percent that did had one thing in common. Someone in finance, not someone in IT, owned the measurement plan from week one.

From the few we have deployed and the dozen more we have audited for clients, the pattern I keep seeing is this. Teams build first, ship, then go looking for a number. By the time they look, the pre-deployment baseline is gone. There is nothing left to compare against. So they pick a flattering baseline, write a flattering number and the project rides on that until the next budget cut, when nobody can defend it.

“ Most AI ROI dashboards are theater. If the number does not collapse when you switch the feature off for half of traffic, you never had a number.

02 Three return layers

Where AI value actually shows up on a P&L

Returns from AI do not arrive in one shape. They arrive in three. Each one needs a different evidence standard. Mix them in a single ROI percentage and you have built a number that cannot be defended in any direction. The McKinsey State of AI survey work going back several waves keeps showing the same split between cost-out, productivity and strategic gains. Most companies report on one, vaguely. Few report on all three with separate methods.

FIG. 01 – RETURN LAYER ANATOMY Direct vs productivity vs strategic

  • Direct cost reduction – Productivity & capacity – Strategic & compounding
  • Difficulty to measure – Low – Medium – High
  • Examples – Cost per support ticket vs human – Reps making more outbound calls – 24/7 retention, conversion lift
  • Measurement window – 1 to 3 mo – 3 to 6 mo – 6 to 18 mo
  • Evidence standard – Point estimate – Holdout group – Range with stated confidence

03 The formula

An ROI calculation that survives a budget review

Formula

ROIY1 = (L1 + L2 + L3) / total investment

L1

Direct savings. Defensible point estimate from invoices and headcount math.

L2

Productivity value. Requires a holdout group or before/after with the same workload mix.

L3

Strategic value. Stated as a range with the confidence interval and the assumptions written down.

investment

Build cost plus Y1 run cost plus a 30% hidden-cost loading for change management, training, prompt and policy maintenance.

One worked example. Klarna announced in February 2024 that its OpenAI-powered assistant handled two thirds of customer service chats in the first month, with a claimed annual profit impact of around $40 million. By 2025 Klarna had publicly walked some of that back and started rehiring human agents for quality reasons. Both things can be true. The L1 cost-out number was real. The L3 strategic claim about a fully automated front line was a hope. A serious ROI write-up would have separated the two from day one.

“ "We expect Year 1 ROI between 180% and 320% with high confidence in the lower bound" beats "Year 1 ROI is 250%" every time you walk into a budget meeting.

Pressure-test your own numbers before that meeting using the interactive AI ROI calculator. It lets you scrub build cost, hidden overhead, run cost and projected value against each other in real time.

04 Coding copilots

Why per-seat tools win the ROI argument almost by default

The cleanest measurable ROI story in AI today is still developer copilots. GitHub published a controlled study in 2022 showing developers using Copilot completed a JavaScript task about 55% faster than the control group, with higher self-reported satisfaction. Later studies from GitHub and academic groups put the effect smaller but still real, often in the 10% to 25% range for realistic enterprise tasks. The reason this calculation survives scrutiny is structural. Per-seat cost is around $20 a month. Fully loaded developer cost is closer to $20K a month. Even a 5% productivity gain that nobody can confidently attribute pays the license back many times over. You do not need a clever ROI model. You need the invoice.

Custom agents are the opposite. They cost real money to build. They ship slower. The value lives in workflows that touch revenue. The math is not harder, but the discipline has to be there from week one. If you are still deciding between buying a copilot and building an agent, read AI agents vs chatbots vs copilots first. The ROI conversation is downstream of that choice.

05 Measurement discipline

What honest measurement actually looks like

Honest ROI discipline
  • Baseline captured before deployment, not retrofitted later
  • Holdout or feature flag in place so you can switch the AI off and measure the gap
  • Success criteria in numbers by week two, signed off by finance
  • Weekly tracking during ramp-up, formal reviews at 90 days, 6 months and 12 months
  • Three numbers reported with ranges and the methodology written underneath
ROI anti-patterns
  • Vanity counts. "10K AI conversations" with no resolution rate or CSAT delta
  • Cherry-picked baseline. The worst pre-AI quarter becomes the comparison
  • AI credited for outcomes it touched but did not drive
  • One quarter post-launch treated as conclusive evidence in either direction
  • No written method. The model exists in one analyst's head and one spreadsheet

06 Knowing when to quit

Graceful shutdown is part of an honest ROI program

Not every AI build pays back. If after six months in production you cannot show movement on the chosen success metric, adoption has stalled under 30% or run cost keeps drifting up with no offsetting revenue, the honest move is to wind it down. The cost of running a stalled implementation is not just the cloud bill. It is the team, the dashboard nobody trusts and the credibility you spend defending it. Shutting down a project that did not work and writing a one-page note on why is the kind of memo that gets you a budget for the next try.

Sunset triggers worth writing into the plan

Six months with no movement on the primary metric. Adoption below 30% with no clear cause to fix. Quality regressions that cannot be closed inside budget. Run cost trending higher than projected without revenue catching up. Decide these triggers at week two, not at month nine. The hidden cost of not using AI article frames the other side of this. Sometimes the bigger risk is doing nothing while peers move.

07 How we would actually approach this

A two-week plan to get an ROI number you can defend

Week one. Pick one workflow. Write down the current cost in time and money with three months of evidence behind it. Decide which of the three layers your project will be measured on, name a primary metric and define a holdout or A/B mechanism before any code ships. Get a finance owner to initial the plan.

Week two. Build the dashboard before the feature. Wire the baseline into it. Schedule the formal review at day 90. Write the sunset triggers and what the next action is in each case. Only then start the build. If the discipline on the first project feels heavy, that is the point. The second and third project inherit the templates and the conversation gets faster every time.

If you want a partner to run this exercise on your first real project, our services page describes how we scope it. Or look at the 90-day implementation roadmap for a longer-form version of the same plan.

Related reading

Frequently Asked Questions

What is a realistic Year 1 ROI for AI implementation?+ For a well-scoped, well-instrumented project we typically see 150% to 400% in Year 1, with per-seat coding copilots clearing that bar at one or two orders of magnitude higher because the unit cost is so low. Custom agents have wider variance. Successful ones return three to ten times. Failed ones return nothing. Customer support is one of the most predictable first projects. The AI customer support playbook includes the outcome benchmarks to anchor a projection. When should we start measuring ROI?+ Before you write any code. Capture the baseline in the same week you scope the project. Begin post-deployment tracking from day one of production. Hold formal reviews at 90 days, 6 months and 12 months. Do not draw firm conclusions before six months in production. Should we use third-party benchmarks or just measure our own?+ Both. Industry benchmarks tell you whether your result is competitive. Internal measurement tells you whether your investment is paying back. When the two disagree, that is usually the most useful signal you have. It is pointing at an implementation problem you need to fix. How do we handle ROI on AI that improves quality but does not reduce cost?+ Quality improvements have a downstream financial effect. A one-point NPS gain in your industry maps to what revenue effect over a year? A half-point retention gain equals what customer lifetime value? Convert the quality metric to dollars with a written conversion method or it will not survive a budget review. Be explicit that this is a modeled number with a range, not an invoice. What discount rate makes sense for multi-year AI investments?+ Most teams use the internal hurdle rate of 10% to 15%. For AI specifically I would lean toward the higher end, 15% to 25%, because of capability decay risk. What is competitive AI today is table stakes in 18 months and the next version costs less to run. A higher discount rate forces the conversation onto what the project does in Year 1 instead of what you hope it does in Year 3. How do we report AI ROI to non-technical leadership?+ Three numbers plus a method paragraph. Total investment, total return, payback timeline. Underneath, one paragraph on what is in the number and what is not, with confidence stated. People trust ROI numbers presented with explicit uncertainty more than precise-sounding numbers presented as fact.

ANM SOLUTIONS / CONTACT US

Need help applying this to your business?

We turn AI insights into measurable business outcomes. Tell us about your workflow and we'll show you the highest-impact place to start.

Related Articles

Implementation

The 90-Day AI Implementation Roadmap: From Scope to Production

A practitioner's view of the first 90 days of an AI implementation. Boring stack, eval-first, kill switch by day 90. Where most plans go wrong and what to do instead.

Read article

Vendor Selection

How to Choose an AI Implementation Partner Without Overpaying

An honest look at the five categories of AI implementation partner, where each one fits and how to tell a useful vendor from an expensive deck factory in a single technical conversation.

AI Adoption

The Hidden Cost of Not Using AI, Honestly Counted

The cost of ignoring AI is real but narrower than the doom pieces claim. A practitioner view on which workflows actually pay an efficiency tax, where the FUD is wrong and how to decide when to move.

Published: 2026-05-05 · Author: A&M Flow