The pruning has started. Most of it is cutting the wrong branch.
Something changed in enterprise AI this year, and it didn't arrive as an announcement. It arrived as a spreadsheet.
In the 2026 State of AI Cost Governance Report, published this month, 62% of organisations said they had hit unexpected AI costs large enough to materially alter a business decision. Forty per cent had to escalate to board level. A third imposed emergency spending freezes. One in four delayed or cancelled an AI initiative outright because of what it cost to run.
Gartner's Anushree Verma put the shift plainly: deployments without tangible value "will now become much harder to sustain." The same firm expects more than 40% of agentic AI projects to be cancelled before the end of 2027.
None of this is the AI bubble popping. It's something more ordinary and more overdue: AI projects being subjected to the scrutiny every other line of capital expenditure has always faced. Pruning, not collapse.
The problem is what the pruning is being done with.
An emergency freeze cuts what's expensive, not what's worthless
Those are different sets, and the overlap between them is smaller than people assume. Some of the most expensive things you're running are the most valuable. Some of the cheapest have been quietly returning nothing for a year and a half.
But cost is knowable by Thursday. You can pull the invoices, add them up, and rank. Value takes a method, a baseline you probably didn't establish, and a willingness to hear that something you championed isn't working. So when the board asks the question and an answer is needed this quarter, cost wins by default — not because anyone decided it was the better criterion, but because it was the available one.
A freeze applied to the expensive things rather than the worthless ones doesn't reduce waste. It reduces spend, and takes some of your best work with it.
That's the failure mode of 2026, and it's quieter than the failure mode of 2024. Nobody writes a post-mortem about the initiative that got frozen in month nine and would have paid back in month fourteen.
Why value is the harder number
Look at how thin the measurement layer actually is.
Most organisations using AI cannot point to any profit impact
Share of all organisations surveyed, by what they can evidence
The largest group by some distance isn't the winners and isn't the failures. It's the organisations where something is genuinely working and nobody can put a number on it. That's a worse position than outright failure, because failure at least resolves. You can stop a thing that visibly doesn't work. You cannot stop a thing that might be working — so it stays funded, and so does the one next to it, and eighteen months later there are nine of them and a budget line nobody wants to defend.
Then the freeze arrives, and every one of those nine is judged on the only axis anyone has: what it costs.
We had this problem ourselves
This isn't a diagnosis we arrived at by reading surveys. We built our first product because we had the problem.
We had more ideas than budget and no consistent way to choose between them. Not a shortage of enthusiasm — an excess of it, unevenly distributed. And in that situation the idea that gets built is often the one whose advocate argued hardest last Tuesday. Not the best idea. The best-defended one, on the day.
The uncomfortable part wasn't the bad decisions. It was that we couldn't tell which ones they'd been. Everything acquires a reason afterwards; the projects that worked have explanations and so do the ones that didn't, and the explanations sound about equally good.
So we built Opptora to be the thing that decides what we build. Every initiative goes through the same questions — what it's actually worth, what it will cost to run, and what would have to be true for it to work — so the answers are comparable with each other rather than with the mood in the room. The mechanism that matters is the boring one: it forces assumptions into the open. You can disagree with an assumption. You cannot argue with a hunch.
Why the costs surprised everyone
The other half of this year's story is that organisations were not wrong to be shocked by the bills. AI cost lands in more places than anyone budgets for, and most of it arrives below the line where anyone is looking.
What gets budgeted, and what actually gets spent
The gap is where 2026's surprises came from
Budgeted
- Licences and per-seat subscriptions
- The build, as a project cost
Actually spent
- Inference — and every retry · 43% name token consumption as their main source of unexpected spend
- Evaluation runs, every time the model underneath changes
- The human review step nobody counted as an AI cost
- The engineer who keeps it alive · not booked against the initiative
- Coding assistants across engineering · Mavvrik 2026: 98% of teams use them; only 42% count them in AI spend
Note the direction of that last number. Forecasting accuracy got worse over a year in which everyone got more experienced. That's what it looks like when spend outruns instrumentation.
What to do instead of freezing
If you're going to prune — and you probably should — prune on value, and accept that establishing value takes a few weeks you'd rather not spend. The alternative is cutting by invoice size and hoping.
The verdict worth reaching on any initiative is one of three:
Three verdicts, and nothing in between
Every initiative gets exactly one, with the evidence behind it
Back it
The evidence holds. Fund it properly rather than keeping it on pilot rations.
Fix it
The case works only if something changes. Name that thing, and set a date to re-check.
Stop funding it
It doesn't pay and won't. Say so early, while the sunk cost is still small.
The method behind them is unglamorous. Establish what's actually deployed, including the things that were never formally launched. Model what it genuinely costs to run — inference and retries, evaluation runs, the review step, the engineer — which is nearly always more than the licence and always more than the budget. Work out what it returns, and where that can't be measured, say so rather than estimate it into existence. Then rank.
Two details do most of the work. Every initiative faces identical questions, which is what makes the verdicts comparable: the point isn't that any single assessment is perfect, it's that they're wrong in the same direction, so the ranking holds. And the instrumentation stays behind afterwards. A verdict is a document, and documents age — what you want is the ability to re-run the question next quarter against what actually happened rather than what was projected.
The part nobody enjoys
Sometimes the answer is stop.
We'd rather say that in week two than year two, and we say it about our own ideas too — Opptora has killed things we were personally attached to, which is the only real test of whether a tool like that is doing its job or just producing agreeable paperwork.
The reason to want that verdict early isn't discipline for its own sake. Every month a dead initiative stays funded is a month the one next to it doesn't get the money. The cost of not knowing isn't the waste. It's the thing you didn't build instead.
If something in your AI portfolio is working and you can't say what it's worth, that's the conversation to have — before someone with a spreadsheet and a deadline decides for you. It's diagnostic, not a pitch.
Sources
- McKinsey, The State of AI: Global Survey 2025 (2025).
- Mavvrik with Benchmarkit, 2026 State of AI Cost Governance Report (August 2026).
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Anushree Verma (press release, June 2025).