Somewhere in every financial services group this quarter, a chief financial officer is staring at the same three artefacts. The first is a stack of business cases — polished, urgent, and unfalsifiable, each promising productivity gains of 30 per cent against a baseline nobody measured. The second is a line item that did not exist two years ago: token consumption, the metered cost of machine intelligence, compounding monthly, which procurement has no template for and no one can forecast to within a factor of three. The third is the revenue line, which has noticed none of it. Even the man selling the tokens has stopped pretending otherwise: Sam Altman of OpenAI conceded last month that customers have exhausted their entire 2026 AI budgets already, and called the return-on-investment question “the most fair criticism right now of AI”. When the vendor validates your unease, your unease is data.
The temptation is to treat this as a technology problem and delegate it. It is not. Clamouring bids, opaque unit costs, unproven returns and a fixed envelope — this is capital allocation under uncertainty, which is to say it is the CFO’s actual job, presenting in unfamiliar clothes. The office that exists to allocate scarce capital against uncertain returns is the natural owner of the AI question, and the sooner it stops policing a budget line and starts constructing a portfolio, the sooner the chaos acquires a shape.
The paradox, named
Begin by refusing to panic about the missing productivity, because its absence is precisely what the economic record predicts. Robert Solow observed in 1987 that the computer age was visible everywhere except in the productivity statistics; the gains took the better part of two decades to arrive, and arrived only for the firms that reorganised work around the technology, not those that merely bought it. Erik Brynjolfsson gave the pattern its modern name — the productivity J-curve: heavy investment in intangibles depresses measured productivity before lifting it. The present data point is brutal and clarifying at once. MIT’s study of enterprise AI last year found 95 per cent of pilots delivering no measurable P&L impact — and diagnosed the cause as organisational, not technological: a learning gap, with budgets piled into sales-and-marketing demos where returns are thinnest, while unglamorous back-office automation quietly pays.
Read correctly, the paradox cuts both ways, and it defines the CFO’s two failure modes. Fund everything on faith, and you are financing the 95 per cent. Freeze everything on the evidence, and you strand the firm on the wrong side of the J-curve while competitors do the reorganising. The posture this column has argued for elsewhere — belief in the destination, audit of every step — has a budget attached, and this is it.
Who to call
The CFO’s telephone now offers four species of caller, each with a tell. The consultancies sell transformation by the hour — and an adviser whose revenue model is hours has a structural conflict when advising on a technology whose entire point is fewer of them. The vendors sell seats and tokens through pricing models that mutate faster than procurement can standardise them; their durability, and the capital cycle financing them, this series has examined sceptically elsewhere, and a CFO should think hard before wiring the operating model into the pricing of a firm whose own economics are unproven. The internal enthusiasts are long on demonstrations and short on baselines. And the internal sceptics — unfashionable observation — are often carrying load-bearing information: their objections are usually a map of where the control gaps sit.
The first calls, however, are none of these. Call the chief risk officer and the head of internal audit, because the control perimeter — what may ship, what may touch regulated data, who answers for machine-speed work — decides what can ever reach production, and the accountability redesign this series has set out elsewhere is the precondition for every pound that follows. Call the chief technology officer, and ask for exactly one number: the fully loaded unit cost of each live use case, tokens included — not the programme budget, the unit cost. If that number cannot be produced within a fortnight, that fact is itself the finding. And make the unfashionable third call: to the head of one well-instrumented operational process — claims, onboarding, reconciliations — because that is where the first honest baseline lives, and the first honest baseline is worth more than the next ten business cases.
The doctrine
Three moves, each recognisably the CFO’s own craft. First, measure before you fund. No case advances without a baseline, and the half-baked decks are converted into a single standard instrument fitting one page: current unit cost of the process; proposed unit cost; token cost at deployment scale, not pilot scale; and a named owner who commits the savings to the budget line rather than to the great euphemism “redeployed capacity”. The pilot-to-scale distinction deserves its moment of arithmetic, because it is where budgets die. A recent analysis of billions of enterprise API calls found firms routing every task to frontier models paying $18.40 per million tokens while firms with a tiered architecture — small models for small tasks, frontier models for frontier tasks — paid $2.31: an 87 per cent cost gap flowing from one architectural decision, made once, at the start, and rarely revisited. A CFO does not need to understand transformers to ask which side of that decision each business case sits on.
Second, run a barbell, not a queue. A small number of deep, production-grade bets in processes the firm controls and has instrumented — where MIT’s 5 per cent actually live — and, at the other end, broad, cheap experimentation under a hard token budget, so the organisation learns fast at bounded cost. Nothing in the expensive middle, which is precisely where the half-baked cases congregate. And kill quickly, publicly, without stigma: in a technology improving this fast, a cheap negative result is genuinely valuable information, and a portfolio that never kills anything is not a portfolio but a queue.
Third — the title’s instruction — govern the unit economics, not the projects. Tokens are a cost of goods sold wearing an operating-expense disguise, and they obey an economics the CFO must internalise now: the price per token is collapsing — down some 98 per cent for frontier-level work since 2023 — while total enterprise spending rose more than 300 per cent last year and has doubled again since, because consumption is growing faster than prices fall. Jevons’s 1865 coal paradox, replayed in silicon: cheaper intelligence is an invitation to consume more of it, and the agentic systems now arriving consume tokens at ten and twenty times the rate of the chatbots they replace, with credible projections of a further twenty-fold rise in consumption by 2030. The pricing page will keep telling you things are getting cheaper. The bill will keep disagreeing. Only per-use-case unit-cost telemetry — built now, before the volumes arrive — tells you whether the firm is buying productivity or renting theatre.
Capacity is not cash
One more discipline, because it answers the question that started this piece: where are the gains? The first returns from this technology do not arrive as revenue. They arrive as capacity — the analyst’s afternoon freed, the backlog cleared, the report produced in an hour. And capacity becomes money only when a manager makes an uncomfortable decision: to take the saving, to redeploy the person, to decline the backfill, to reprice the service. Absent that decision, the gain is real, invisible, and silently absorbed — which is the true story behind most of the missing productivity. The gains are not absent. They are unharvested. That is not a technology failure; it is an operating-model failure, and it lands on the CFO’s desk, because only the budget can compel the harvest: the savings are claimed at the moment of funding, or they are never claimed at all.
The bill you have seen before
None of this should feel novel, because the CFO has lived it once already. Cloud computing performed the same trick a decade ago — capital expenditure dissolved into consumption pricing, the bill shock arrived on schedule, and the discipline for governing it, FinOps, was invented years too late and at great expense. The sequel is running at speed: the share of those practitioners now handed responsibility for AI spend leapt from roughly a third to nearly all of them within a single year. The lesson of the first film is simple — build the discipline before the bill this time, not after. And the deeper counsel is the one this column keeps returning to. The CFO cannot make the technology work; that was never the job. What the CFO alone can do is make the organisation legible to itself — baselines, unit costs, owners, harvested savings — so that when the fog thins, the firm knows precisely where it stands. In conditions like these, the firm that can measure is the firm that can move. Follow the tokens. They are the only witnesses that never exaggerate.