← Blog
StrategyJun 24, 2026· 10 min read

How to Evaluate AI ROI — and When to Kill the Pilot

Most AI pilots don't fail loudly. They drift. Here is the scorecard executives use to decide, and to kill on time.

Key takeaways
  • AI pilots rarely fail with a clear signal — they drift into permanent 'promising' status, which is the actual failure mode to design against.
  • ROI evaluation for AI needs a different denominator than traditional software: include labor re-absorption and model drift cost, not just license spend.
  • A pre-committed kill date and kill criteria, set before launch, is the single highest-leverage governance step most companies skip.
  • Sunk-cost bias is stronger with AI pilots because they feel novel and strategic, which makes an outside scorecard more necessary, not less.

Ask any operating executive how many AI pilots their company has running, and you will usually get a confident number. Ask how many of those pilots have an actual end date and a kill criterion written down before launch, and the confidence disappears. That gap — pilots without a defined exit — is the single largest driver of wasted AI spend in large organizations today, more than model cost, more than integration overhead, more than talent.

The reason is structural, not a failure of any one team. Traditional software pilots fail loudly: the vendor misses a deadline, the integration breaks, the users refuse to log in. AI pilots fail quietly. The model mostly works. Accuracy sits at 'pretty good.' Someone on the team likes it. Nobody wants to be the person who cancels the innovative project, so it rolls forward another quarter, and another, consuming budget and attention while never quite proving itself and never quite dying.

Build the ROI denominator correctly first

Most AI ROI calculations undercount cost because they only capture the license or API bill. A defensible calculation includes at least four cost categories, three of which are usually missing from the first draft any team presents to an executive.

Cost categoryCommonly missed?Where it hides
Model / API spendNo — usually capturedFinance system, easy to find
Integration & maintenance engineering timeYesCharged to general engineering overhead, not the project
Human review / correction laborYesAbsorbed into existing job descriptions, never timed
Drift monitoring and re-tuningYesAssumed to be a one-time cost; it is not

The human-review line is the one that most changes the picture. An AI system that produces a draft a person must still fully verify is not automating the task — it is changing its shape. That can still be worth doing, but only if the calculation is honest about the labor still required, measured in hours, not vibes.

Set the kill criteria before the launch, not after

The organizations that kill pilots on time share one habit: the kill criteria are written into the pilot charter before a single user touches the system, and they are specific enough that no one in the room can argue with the outcome later.

  • A hard end date for the evaluation window — typically 60 to 120 days, not 'ongoing.'
  • A quantified success threshold tied to a business metric the pilot was funded to move, not a proxy metric like 'usage' or 'satisfaction.'
  • A named decision-maker who commits, in writing, to make the call on the end date — not a committee that can defer.
  • An explicit statement that 'promising but inconclusive' at the end date defaults to kill, not extend, unless the decision-maker actively overrides it with a new, dated success threshold.
The default outcome for an AI pilot without a written kill date is not success or failure. It is permanent limbo, which is more expensive than either.

The four questions that replace a vague check-in

Rather than a status meeting that asks 'how's the pilot going,' the scorecard reviews four specific questions, in this order, because the order matters — later questions are irrelevant if earlier ones fail.

  1. 1Did the pilot move the business metric it was funded to move, by the threshold set in advance?
  2. 2Is the fully loaded cost per outcome — including review labor — lower than the process it replaced, or trending there on a known curve?
  3. 3Has accuracy or output quality been stable across the evaluation window, or is it drifting in a direction that needs re-tuning we haven't budgeted?
  4. 4If we scaled this from a pilot group to the full function, does the cost and oversight model still hold, or does it only work at small scale?

Question four is where a surprising number of otherwise-successful pilots actually die, and rightly so. A pilot run by five enthusiastic early adopters who review every output carefully often cannot survive contact with five hundred users who will not.

Why sunk-cost bias hits AI pilots harder

Executives are generally disciplined about killing underperforming initiatives. AI pilots get more protection than they deserve for a specific reason: they carry a narrative of strategic necessity that a failing CRM rollout does not. Killing an AI pilot can feel like admitting the company is 'behind,' so it survives past its natural evaluation point on reputational inertia rather than performance. The scorecard exists precisely to remove that emotional variable by making the kill decision mechanical rather than a judgment call made under social pressure.

A framework built for the decision, not the demo

The ROI and portfolio-management module in AI Executive Mastery gives C-suite participants the exact scorecard and kill-date template used above, applied live to a pilot from their own business during the course — so the first time they use it isn't under pressure in front of the board.

What a healthy AI pilot portfolio looks like

A mature portfolio kills a meaningful share of its pilots — often a third or more — on schedule, without drama, because the criteria were set in advance and everyone agreed to them before anyone got attached to the outcome. That kill rate is not a sign of poor selection. It is a sign the evaluation process is actually working, in the same way that a venture portfolio with a zero failure rate is a portfolio that took no real risk. The goal was never to make every pilot succeed. It was to find out fast, cheaply, and honestly which ones would — and to stop funding the ones that would not, on the date you said you would.