AI use cases: measuring the value
Every AI project ends with a claim. “It saves us hours.” “The team is so much faster now.” The claim is free; anyone can make it, and after months of effort everyone wants it to be true. Verified savings is a formula, not a word, and the formula has a strict input order: baseline first, build second, measurement third. Reverse the order and no amount of goodwill can recover the proof.
Freeze the baseline before you start
Section titled “Freeze the baseline before you start”The baseline is the unit cost of the process as it runs today: minutes per invoice, cost per resolved ticket, days per report. It has to be pulled from a system before the build begins, because after go-live the old process no longer exists to be measured. A baseline reconstructed afterwards from memory is an estimate wearing a lab coat.
A frozen baseline is specific and written down: which counter, from which system, over which window, pulled on which date.
Weak: "Support says a ticket takes about half an hour."Frozen: "Jan to Mar: 1,240 tickets resolved, 26 min average handling time, from the ticketing system's own reports, exported 2026-04-02."If the number can’t be produced now, it won’t be producible when it’s time to prove value. That’s the measurement gate from the three gates, applied with a deadline: the export exists before the first line of the build.
The formula
Section titled “The formula”Once the baseline is frozen, verified savings is arithmetic:
verified savings = (baseline unit cost - new unit cost) x actual volume - run costs (tokens, licences, hosting, upkeep) - governance costs (review time, audits, monitoring) - rework costs (fixing what the system got wrong)Three things keep this honest.
Actual volume, not projected volume. The savings accrue on the units that really flowed through the new process, read from the same counter as the baseline. A forecast is part of the business case; it has no place in the verification.
Net, not gross. The gross difference is the number that gets quoted in slides. The net number is the one that survives an audit:
| Cost line | What goes in it |
|---|---|
| Run | Token or licence spend, hosting, the integration someone now maintains |
| Governance | Human review at your chosen autonomy tier, sampled audits, monitoring, incident handling |
| Rework | Time spent correcting wrong outputs downstream |
Rework is not a hypothetical line: a sizeable share of AI-assisted commits ship a defect (see The deterministic layer for the numbers). Whatever your domain’s equivalent of a bad commit is, budget for catching and fixing it, and subtract that time from the win.
Your numbers, not anyone else’s multipliers. A vendor’s “3x productivity” is their testimony about someone else’s process. The formula only takes inputs you measured yourself.
Measure the process, not the people
Section titled “Measure the process, not the people”Put the measurement on the process counter: tickets resolved, invoices posted, cycle time from intake to decision. Not keystrokes, not screen time, not per-employee dashboards.
This is partly pragmatic. Process counters are where the money actually moves, and they already live in systems you can export from. But it’s also protective: the moment measurement lands on individuals, it becomes performance surveillance. That poisons the cooperation you need from the very people whose workflow is changing, and it raises real questions under data protection and employment law.
Without a baseline, it’s testimony
Section titled “Without a baseline, it’s testimony”Why insist on all this? Because perceived improvement and measured improvement routinely disagree, in both directions.
In a 2025 randomized trial by METR, experienced open-source developers expected AI tools to make them 24% faster, and still believed they had been sped up after the fact. Measured against the clock, they were 19% slower on those tasks. Sincere, expert, first-hand testimony, and wrong. Not because anyone lied, but because “it felt faster” is not a measurement.
The counterexample proves the same rule. Brynjolfsson, Li and Raymond studied a generative AI assistant at a customer support operation (NBER Working Paper 31161) and found productivity up 14% on average and 34% for the newest agents. That claim holds up precisely because the setup matched the formula: a hard process counter (issues resolved per hour), a baseline recorded before rollout, and actual volume. It also shows the gains were wildly uneven across groups, which is one more reason a headline multiplier from someone else’s study says little about your process.
An improvement claim without a frozen baseline is testimony. It may even be true. You just can’t know, and neither can the person who has to sign off on the next one. This is the same standard as the deterministic layer applied to money: evidence, not assertion, and the check defined before the work starts.
Where to go next
Section titled “Where to go next”The formula is the last screen before the decision. The capstone puts all of them in sequence: