A procurement question reaches most finance functions sooner or later, phrased with apparent precision: which model should we standardise on? It sounds like the kind of question that has an answer — a single best system, identifiable from the published rankings, to be chosen and rolled out. It does not, and the belief that it does is among the more expensive misconceptions in enterprise AI.
The reason is that the same model, left entirely unchanged, can succeed or fail at the same task depending on what surrounds it. Give a frontier model a coding problem cold and it may solve a modest fraction; place the identical model inside a well-built scaffold — one that lets it read the codebase, run the tests, see its own errors and try again — and its success rate on the very same problems can multiply severalfold. Nothing about the weights changed. What changed was the harness. Capability, then, is not a fixed property of a model, to be looked up. It is a property of an assembly, of which the model is only one part — and a firm that chooses on the first part alone is deciding on a third of the evidence.
The assembly, not the component
The first part is the model: the trained weights, the raw aptitude. This is what the headlines rank and what most buyers believe they are choosing. It matters, but it is the component, not the machine.
The second is the harness — the software the model runs inside. A racehorse in harness is only as good as the rig it pulls and the driver behind it, and the machine-intelligence sense of the term is no different. The harness is the prompt that frames the task, the tools the model can reach for, the documents retrieved and set before it, the memory it carries between steps, and the control logic that lets it loop, check its work and recover from a mistake. Two firms running the identical model behind different harnesses do not have the same capability; they have different machines that happen to share an engine.
The third is run-length — how long, and how hard, the model is allowed to work before it must answer. The most consequential shift of the past two years is that thinking time has become a capability dial. A model permitted to reason at length, to draft and revise, or to make several attempts and select the best, will clear problems that the same model, replying in one breath, cannot. You can, within limits, buy capability by the minute. And the three are not independent: a heavier harness and a longer run can lift a cheaper model over a task that a dearer one clears unaided. Which is the first reason the ranking cannot answer the procurement question — it holds two of the three variables fixed, and hides them.
Why the leaderboard cannot tell you
That is not the only reason to distrust it. The public benchmarks — the standardised tests on which models are scored and marketed — are useful for tracking the frontier, but each measures a thin slice of what a model can do, and each carries a familiar pathology. Some are saturated: the best models now sit so near the ceiling that the scores no longer separate them. Some are contaminated: the questions, being public, seep into the next generation’s training data, so a high mark may measure memorisation rather than reasoning. And all are subject to the oldest law in management — once a measure becomes a target, it stops being a good measure. A model can be tuned to shine on the tests everyone quotes.
The deeper problem, though, is not that benchmarks are gamed. It is that they answer someone else’s question. A leaderboard measures performance on a generic task, under a fixed harness, against data you did not choose. Your task is not that task. The reconciliation of your ledgers, the triage of your complaints, the screening of your counterparties — each has its own data, its own edge cases and its own definition of a right answer, and no public score speaks to them. The constructive conclusion follows directly, and it is the single most useful thing to take from all this: the only evaluation that counts is your own — your representative data, your acceptance criteria, run against the whole assembly you actually intend to deploy. Build that evaluation once and it will tell you what a thousand leaderboards cannot. Read the leaderboard instead, and you have measured a stranger.
The cost of being right
There remains the figure every CFO reaches for first, and it is the one most often measured wrongly. Token cost — the price of the words in and out — varies not by a little but by orders of magnitude across the available models. It is tempting to read that spread straight off the price list. It is also a trap, because the per-token price is not the cost of the work. A cheaper model may need a heavier harness, more retrieval, longer reasoning and several attempts to reach an answer you can accept, consuming many times the tokens as it goes; a dearer model that succeeds cleanly, first time, can cost less to finish the job. The number that matters is therefore neither the price per token nor the price per call, but the cost per successfully-completed task — the total spend, across the whole assembly, to produce an output that meets your bar. Measured that way, the league table by price can invert: the expensive model is sometimes the economical one, and the cheap model the extravagance. A finance function is well equipped to think in exactly these terms, because it is unit economics applied to inference — cost per good outcome, not cost per unit of raw input.
What a regulated firm should optimise for
Two conclusions complete the picture, and both weigh more heavily in a regulated institution than anywhere else.
The first is that the weighting of the three factors is itself a property of the task, which is why there is genuinely no universal answer. For agentic, multi-step work — an assistant that plans, calls tools and carries a task across many turns — the harness and the run-length dominate, and the underlying model, while not irrelevant, is one input among several. For simple, single-shot work — lifting a figure from a document, classifying a message, returning a short answer — the harness is thin and the model dominates. The shape of the right answer changes with the use, and a firm that standardises on one model for everything will have overpaid for the easy jobs and underbuilt for the hard ones. Horses for courses is not a platitude here; it is the operative principle.
The second is the sting in the tail, and it cuts against the grain of the whole enterprise. Every lever that buys more capability — the elaborate harness, the long agentic run, the many attempts — also makes the system harder to govern. It grows less deterministic: the same input, run twice, need not return the same output. It grows more variable in cost: an agent free to loop and retry can quietly consume a budget. And it acquires failure modes that belong to the harness rather than the model — a tool invoked wrongly, a retrieved document that misleads, an instruction smuggled in through the data. For a regulated firm, capability is not the only axis and often not the decisive one. The best system for a decision that bears on a customer or a balance sheet is frequently not the most capable per pound, but the one whose behaviour can be bounded — its cost capped, its variance constrained, its failure surface small enough to inventory and to test.
Here the discipline of the previous era reasserts itself. Model-risk management — the PRA’s SS1/23, the Federal Reserve’s SR 11-7 — requires that anything bearing on a material decision be inventoried, validated, monitored and owned. What this thesis adds is that the object of that governance is not the model but the assembly: the harness is in scope, the run-length is in scope, and the cost envelope, the determinism and the failure modes of the scaffold must be validated as one system, because it is as one system that it will succeed or fail in production. An earlier piece held that openness of weights is not auditability of a model. Its companion belongs beside it: capability of a model is not fitness of a system.
The honest answer to “which model is best”, then, is another question: best at what, wrapped in what, allowed to run for how long, and measured against whose standard? That is not evasion; it is the discipline the question demands. Define the task; build your own evaluation; measure the cost of a completed job, not the price of a token; weight the model, the harness and the run-length to the work in front of you; and, for anything that touches a regulated outcome, prefer the assembly you can bound to the one that tops a chart. There is no horse that wins every course — only the right horse, in the right harness, over the right distance, and the sense to know which race is being run.