This isn't a cost-cutting story. It's a rewiring story. New line items appear. Old ones invert. Margin structures that looked like laws turn out to have been artefacts of a labour-priced business.
Most organisations will meet this late, through pricing pressure they didn't see coming. A few will meet it early, on purpose. Here is what changes, and what your leadership team needs to ask before the next planning cycle.
01. The agentic shift rewrites the ledger.
AI used to buy you a better interface. Now it buys you execution.
For the first years of enterprise AI, spend bought conversation. A person asked, a model answered, the person did the work. Cost stayed exactly where it had always been: in salaries.
Agentic AI moves the work itself. An agent breaks down a goal, calls tools, writes and runs code, checks its own output, retries and closes the loop. The person specifies and verifies. The marginal unit of production stops being an hour. It becomes a token.
The benchmark that matters has moved with it. Not knowledge, reasoning or exam scores, but how long a task an agent can complete on its own, reliably. It is the only AI benchmark with a line on the P&L, and it is rising fast.
The cost curve underneath is more interesting than the headlines. The price of a unit of model output keeps falling. Total AI spend keeps rising, and it regularly comes in above plan. Both are true for a simple reason: agentic workflows consume many times the tokens of a chatbot query, so an outcome delivered by an agent can cost far more in compute than the question it replaced.
That is the central financial fact of this era. Deflation per unit. Inflation per outcome. And a compute bill that behaves like cost of goods sold, not IT. Most AI compute is now inference: recurring operating cost, not lumpy capital spend. Your gross margin is now set by model routing decisions made by engineers.
02. The death of legacy departmental benchmarking.
Every industry has a signature ratio. Every one of them measures the same thing: the cost of human throughput per unit of output. Different clothing, same number.
- Banking has cost-to-income. A decade of digital investment went into the front end; the back office stayed fragmented and staffed.
- Insurance has the expense ratio, moved a point at a time over years of effort.
- Professional services has the billable hour, and much of the work behind it is automatable.
- Healthcare has cost per administrative head.
- Retail has sales per employee and labour as a share of sales.
- Manufacturing has overhead recovery per direct labour hour.
- Software has the split between sales and marketing, R&D and G&A.
None of these were laws. They described how much human labour a function needed to move a unit of revenue, and they held steady for long enough that leadership teams mistook them for physics.
The evidence that they are breaking is commercial, not theoretical. Carriers that run AI with discipline are recovering combined-ratio points faster than the sector managed in the previous decade. The sharpest signal comes from legal, because there the client did the maths first: corporate legal departments have started writing into billing guidelines that the cost of AI-generated work will not be passed on.
The benchmark didn't collapse from the inside. The buyer refused to fund it.
Expect that pattern everywhere. Your customers will reprice the labour in what you sell before your finance team does.
Three benchmarks to retire this quarter.
- Revenue per employee. AI-native companies run at multiples of what incumbents consider top quartile. Nudging the figure up and calling it progress is optimising the wrong denominator. The question is revenue per unit of orchestrated capacity, people and agents combined. That is the capacity you now buy.
- Any budget line anchored to a headcount or a seat. Software vendors are moving off per-seat pricing to usage and to price per unit of work (per resolution, per qualified lead, per resolved conversation), and they price agents against a salary, not a licence. Your vendor spend is becoming a variable cost of production. Treat it as a fixed pool negotiated once a year and you will be wrong twice a year.
- Every unit cost that prices a person doing a task. Cost per ticket, per claim, per invoice, per case, per line of code. All of them assume a person does the work, so none of them can tell you whether a process got cheaper or got replaced.
The replacement set, and who owns it.
Legacy benchmarking survives because nobody owns the alternative. Assign these before the next planning cycle.
| metric | what it measures | owner |
|---|---|---|
| Cost per outcome | The all-in cost of a completed unit of work, including compute, escalation and rework. | CFO |
| Execution accuracy | First-pass completion without human correction. | COO |
| Rework rate | The hidden line that kills naive agent business cases. | COO |
| Compute efficiency | Useful output per pound of inference, per workflow. | CIO / CTO |
| Orchestration ratio | Agents productively supervised per person. | CHRO with COO |
| Capability half-life | How fast a critical skill or model advantage decays. | CHRO |
| Exposure per autonomous decision | Value at risk where an agent acts without review. | CRO / general counsel |
Two of those sit with functions that have never held a productivity metric. That's the point. This isn't a finance exercise with operational consequences. It's an operating model change with financial symptoms, and the symptoms arrive last.
03. The re-imagined departmental spend.
Research and development: from headcount to orchestration.
Engineering budgets stop being a function of who you can hire. They become a function of how much verified output you can supervise. A growing share of new code is AI-generated and engineer-approved, and teams that use AI daily merge more work.
R&D spend splits: fewer, more senior salaries (the people who can specify systems and judge output) plus a metered compute line that scales with ambition rather than hiring capacity. And the balance between sales and marketing and R&D inverts. For twenty years organisations spent more acquiring customers than building product, because building was scarce. When building becomes cheap, product stops being the moat.
Sales and marketing: content goes to zero, trust gets expensive.
Content production cost collapses towards the price of inference. Content stops being a differentiator and becomes an input. What gets expensive is everything AI can't manufacture:
- Distribution. Owned audiences, channel positions and integration real estate. Infinite content raises the price of attention.
- Trust verification. Provenance, proof, third-party validation and human reference. In a market full of synthetic claims, verified ones command a premium.
- Proprietary data. The one input agents can't generate. Acquiring, instrumenting and defending it becomes strategy, not a marketing tactic.
Operations and G&A: headcount out, fixed compute in.
This is where the change is bluntest. Agents handle most routine queries with no human involvement, and high-volume, high-repetition roles compress sharply, replaced by a compute line that behaves like infrastructure: largely fixed, forecastable and cheap at the margin.
The people who remain are not junior. They handle exceptions, escalation and quality assurance on agent behaviour, and they are paid for it. Service becomes a small senior team plus a metered system, measured in tokens per resolution and containment rate rather than average handle time. G&A follows with a lag: every function you automate adds a governance obligation that didn't exist before.
04. Flat organisations, dynamic teams.
Hierarchy exists to solve an information problem. Software now solves the same problem.
People have a limited span of control, so organisations built layers to delegate, aggregate and report. Orchestration software does that in code, continuously. Large employers have started to frame cuts explicitly around fewer management layers.
The emerging unit is the conductor: one senior person owns an outcome and directs a set of specialised agents against it. Not a manager of people. An owner of a workflow, accountable for output quality rather than team size.
There is a cost the enthusiasts miss. Middle management is where organisations grow future leaders and hold institutional context. Cut it hard and you create a succession problem that surfaces years later, a liability that never appears on the P&L. Flattening is a financial win and an organisational risk at the same time. Price both.
Team structures stop being annual, too. Capacity is elastic, so teams form around problems and dissolve when they are solved. The org chart stops working as a budgeting instrument.
05. The insourcing paradox.
Labour arbitrage can't compete with capability arbitrage. The traditional outsourcing model is being dismantled as voice and text agents match people on routine interactions. But the conclusion most leadership teams draw, to bring everything back in house, is wrong. The decision splits along a new line.
Insource execution. Anything touching proprietary data, customer relationships or regulated process moves inside, run on models under your control. Not because it is cheaper in isolation, but because execution on your data is where your advantage compounds. Every agent run leaves a record of what worked, what failed and which exceptions recurred. Outsource execution and you hand that asset to a vendor.
Outsource the algorithmic layer. Frontier models, training runs, inference capacity at scale, evaluation harnesses and orchestration platforms are capital-intensive and go out of date fast. Building them yourself is the modern equivalent of generating your own electricity.
Rent the intelligence. Own the context.
Get it backwards, building your own models while outsourcing the workflows that generate proprietary data, and you will have spent heavily to end up commoditised.
06. New line items on the radar.
Four costs that don't exist in most charts of accounts today. Within two planning cycles they are material.
- Model depreciation and fine-tuning. A fine-tuned model is a depreciating asset with an unusually short life: the base model is superseded, the data drifts, the tuning decays. It needs an amortisation policy, a refresh budget and a re-evaluation line. Most organisations have none of the three.
- Data provenance and IP liability insurance. Indemnity on AI output is now a procurement variable, not a legal footnote, and a dedicated AI liability insurance market has formed. Watch the gap: contractual IP indemnity you give customers may not be covered by your existing technology errors and omissions policy.
- Compute elasticity buffers. Agentic workloads are spiky and self-amplifying. An agent that retries, decomposes further or escalates multiplies consumption with no change in demand. Provision a buffer with hard circuit breakers and treat it like a treasury reserve.
- Agent supervision and evaluation. The line nobody budgets: continuous evaluation, red-teaming, human review and rework. It is the whole difference between the unit economics in the business case and the ones you actually get.
07. The P&L shift.
Two organisations with identical revenue: one built for the labour-priced world, one rewired for agentic execution. Where each line moves:
| line item | as a share of revenue | what drives the cost | the new benchmark |
|---|---|---|---|
| Cost of goods: human delivery and service | Falls | Senior exception handlers only | Contained resolutions per human touch |
| Cost of goods: compute, hosting and inference | Rises | Tokens consumed per delivered outcome | Gross margin per million tokens |
| R&D: human engineering | Falls | Senior specification and verification | Output surviving in production per engineer |
| R&D: agent compute and orchestration | New | Autonomous build and test cycles | Work shipped per token |
| Sales and marketing | Falls | Distribution and trust, not content | Pipeline per pound of distribution spend |
| Customer service and support | Falls | Escalation and quality assurance on agent behaviour | Tokens per resolution, containment rate |
| General and administrative | Falls | Controls and governance, not processing | Close cycle time per finance FTE |
| Model depreciation and fine-tuning | New | Refresh cadence and data drift | Performance retained per pound of refresh |
| Data provenance and IP liability insurance | New | Output exposure and indemnity gaps | Insured exposure against AI-touched revenue |
| Compute elasticity buffer | New | Workload spikes and agent retry loops | Buffer use against breach events |
The first read shows margin expanding. The second shows what matters: a meaningful share of revenue moves into cost categories with no owner, no policy and no benchmark. That is the exposure.
Note what didn't happen. Revenue is identical. This is a structural cost story, not a growth story, which means it is available to incumbents, not just AI-natives, and it gets competed away in pricing faster than most boards expect. The margin doesn't stay with you. It becomes the price of entry.
The warning.
Benchmark against a pre-AI peer set and you are optimising your position in a league that is being dissolved.
The dangerous position isn't being behind on AI. It's being on benchmark: comfortable at median spend, median revenue per head and median cost per transaction, while a competitor runs the same revenue on far fewer people and reinvests the difference in distribution and data you can't match. Legacy metrics will keep telling you you're fine right up until the pricing pressure lands.
Three questions for your leadership team this quarter.
- What is our all-in cost per completed outcome in our three highest-volume workflows, including compute, escalation and rework, and who owns that number? If nobody owns it, you can't see your own unit economics, and every agentic business case you approve is a guess.
- What is our total AI spend, split fixed and variable, and what happens if agent consumption trebles with no change in demand? Model the spike, set the circuit breakers and provision the buffer before you need it.
- Which contracts assume IP indemnity we don't carry insurance for, and how much AI-touched revenue sits behind that gap? This is the one most likely to become a board problem with no warning.
The organisations that come through this well won't be the fastest adopters. They will be the ones that rewired their measurement system before they rewired their operating model, because you can't manage a cost structure you can't see. Your chart of accounts is a strategy document. Update it.
Related: how to measure the ROI of AI, rewiring the operating model for AI, and the AI ROI hub.