EXECUTIVE SUMMARY - AI business cases fail in finance review for one reason: they present a number where the CFO needs a model. A number is an assertion; a model is an argument - inputs a finance team can interrogate, assumptions with named owners, and sensitivity ranges that admit uncertainty instead of hiding it. This report introduces the CFO-Grade Value Stack - three layers of AI value that must be modeled differently - and the Five Defensibility Tests I apply to every value-engineering engagement before a number is allowed into an executive deck. The uncomfortable rule underneath both: if value can't be defended in front of a CFO, it isn't value yet.
Why AI cases die in finance review
The failure is rarely arithmetic. It is category confusion: hard savings, soft capacity gains, and strategic option value get summed into one headline figure - "this delivers $4.2M" - and the first probing question collapses the whole structure, because the presenter cannot say which dollars are bankable and which are hopeful. Finance teams do not punish uncertainty; they punish uncertainty disguised as precision.
The CFO-Grade Value Stack
Layer 1 - Cost takeout (bankable). Spend that stops: license retirements, vendor consolidation, error and rework costs eliminated, overtime removed. Evidence standard: current invoices and incident logs. Model at face value with a timing lag. This layer alone should justify the project's floor case.
Layer 2 - Capacity release (conditional). Hours returned to people - the largest and most abused layer. Hours saved are only worth money if they convert to something: headcount avoidance, throughput, backlog reduction. Evidence standard: a named conversion mechanism per line. Model at a haircut (I default to 50–70% realization) with the mechanism written next to the number. An "efficiency gain" with no conversion story is a morale benefit, priced at zero.
Layer 3 - Optionality (strategic). Revenue upside, faster launches, decisions made earlier. Real, but unbankable ex ante. Evidence standard: comparable precedent or pilot signal. Model separately, never summed into the headline - presented as "and if it works, here is the upside case." CFOs respect an upside case clearly labeled as one; they distrust it laundered into the base case.
The Five Defensibility Tests
Before any figure enters a deck, it must pass:
1. The Source Test. Every input traces to an invoice, a log, a timestamp, or a named person's estimate - labeled as such. "Industry benchmark" without a citation fails.
2. The Owner Test. Each assumption has a human owner who will defend it in the room. Orphan assumptions are where models go to die.
3. The Sensitivity Test. The model shows what happens at 70% and 130% of key assumptions. A case that only works at 100% of everything is not a case; it is a bet dressed as one. One tornado view beats ten pages of prose.
4. The Symmetry Test. Costs get the same optimism as benefits. If benefits ramp over three years, implementation pain and run-rate costs don't magically land in month one at list price and then vanish. Include internal time, change management, and monitoring headcount - the lines vendors omit.
5. The Payback Test. Lead with payback period, not five-year NPV. NPV answers "is this worth doing?"; payback answers "how long am I exposed?" - and exposure is the question executives are actually asking. Sub-24-month payback with a defensible floor case closes; a giant NPV with 4-year payback stalls.
Sensitivity is the argument, not the appendix
The strongest thing a value model can do is show the conditions under which the project fails. A sensitivity grid that reveals "NPV stays positive down to 60% benefit realization" is more persuasive than any headline number, because it demonstrates the presenter has already tried to kill their own case. In my deal work, the sensitivity page consistently generates more CFO engagement than the summary page - it is where trust is built.
FAQ
Should I include productivity gains at all? Yes - in Layer 2, haircut, with a conversion mechanism named per line. Exclude them entirely and you undersell; include them raw and you get discounted on everything.
Point estimate or range in the headline? Headline the floor case (Layer 1 + haircut Layer 2), show the range beneath it. Anchoring low and beating it builds multi-deal credibility; the reverse spends it.
What discount rate? The client's WACC if they'll share it; 10–12% as a stated assumption if not - with rate sensitivity shown so the choice is visibly not doing the work.
Action checklist
- [ ] Re-sort your current business case into the three layers
- [ ] Price every Layer-2 hour with a named conversion mechanism - or haircut it to zero
- [ ] Move all Layer-3 upside out of the headline into a labeled upside case
- [ ] Run the Five Tests; fix every fail before the next review
- [ ] Add a sensitivity grid; lead the finance conversation with it
- [ ] Restate the headline as payback period first, NPV second
Related: The Four Gates · Download the working Enterprise AI ROI Model (XLSX) this article describes.
Building a case for a live deal? Value-engineering sprints - model, sensitivity, and the deck - are my core engagement. → Book a consultation
Working through this in your organization?
I advise enterprise teams on exactly these problems. Start with a free 30-minute consultation.
Discuss Your Challenge