We see the same pattern a lot. Licences get bought, a couple of pilots run, and twelve months later nobody can say what changed. The money usually went on tools before anyone agreed what they were meant to move.
Where we do see a return, one person owns one number, and it was measured before anything was built. That part is dull. It is also the part that pays. The method below is how we set it up.
Eight steps to a return you can prove.
- Name the number before you build. Write down the line in the P&L the use case is meant to move (hours spent matching invoices, cost per claim, error rate in a report, conversion on a quote) before any money is spent. If nobody can name one, it is an experiment, and it should be funded and judged as one.
- Set the baseline. Measure that number as it is today, over a period long enough to smooth out the normal ups and downs. Without a baseline there is nothing to measure the result against, and every later claim is an estimate.
- Give it one owner. One named person in the business, not the AI team, owns the number and is accountable for it moving. They sign off the target and report the result.
- Agree the target and the date. How far the number should move, and by when. Agree it with finance, so the result will be believed when it arrives.
- Count the full cost. Licences and seats, compute and inference, integration and data work, change and training, and the costs people forget: human review of the output, rework when it is wrong, and the evaluation needed to keep it right. The all-in cost per completed outcome is the figure that matters.
- Measure against the baseline, in the P&L. Report the change in the named number, in cost, time, quality or revenue. Where you can, compare against a team or period that didn't change, so seasonality or a price rise isn't counted as AI value.
- Keep leading indicators separate. Adoption, usage and user satisfaction tell you whether the return is coming. They are not the return. Report them, but never add them up as value.
- Decide: scale, fix or stop. At the agreed date, scale what moved the number, fix what is close, and stop what didn't. Redirect the budget to the next use case with a clear owner and baseline.
Stop reporting activity as value.
Most AI dashboards report the left-hand column. Boards should ask for the right-hand one.
| what often gets reported | what shows a return |
|---|---|
| Licences bought and seats active | Hours or cost removed from a named process |
| Pilots launched | Use cases live in core workflows |
| Prompts run and usage growth | Error rate, cycle time or revenue moved against a baseline |
| Demo results | Value reported in the P&L |
| Time saved per person, self-reported | Capacity redeployed or cost removed, confirmed by finance |
The calculation.
Return on investment is value realised minus full cost, divided by full cost. The arithmetic is simple. The discipline is in both halves: value realised is the change in the named number against its baseline, confirmed by finance; full cost is everything it took to get there, including the supervision and rework that business cases leave out.
For agentic work, track cost per completed outcome alongside it. Compute cost per unit falls, but agents use far more of it per task than a chatbot, so the cost of an outcome can rise while the price of a token falls. The post-AI P&L sets out why.
Where this sits in an engagement.
The baseline and the owner are set in the first review, when each use case is sized in pounds. Measurement runs through delivery, and the value management office keeps tracking every live use case against its baseline after we have gone. It is the define and realise of the Lumo value framework.
Related: where is the return on AI? (the hub, kept current with source-checked evidence) and rewiring the operating model for AI.