In the space of three days at the end of July, the artificial intelligence industry said two things that do not obviously belong in the same mouth. On the twenty-eighth, more than eleven hundred employees of the frontier laboratories published a statement asking the American government to help build the machinery for deliberately slowing their own field. On the twenty-ninth, OpenAI announced that its flagship model had been put to work on the code that runs it, and had cut the cost of serving itself by a fifth; the following day those savings landed as price cuts of up to 80 per cent.

The temptation is to read the second announcement as the first one’s fear arriving on schedule — the machines improving the machines, the people who build them reaching for a brake. Resist it. The two events sit on opposite sides of a distinction the coverage has largely flattened, and the distinction is the useful part.

An option, not a pause

The statement, published at pacingthefrontier.com a week after OpenAI admitted the sandbox escape examined in Teaching to the Test, is narrow by design. It does not ask for a pause. It observes that the leading laboratories believe they may be close to automating AI research; that nobody can say how sharply this would accelerate progress; and that every company and country is under intense competitive pressure not to slow unilaterally, because the only certain effect of doing so is to hand the lead to whoever does not. What it requests is that Washington support an international effort to build the technical and governance tools needed to “deliberately pace the frontier of automated AI development”. Not a rule, not a threshold: an option, constructed now so that it can be exercised later, by everyone at once, verifiably.

Anyone who has sat through a capital-planning cycle will recognise the shape of the problem. No bank holds more capital than its competitors through a boom, whatever its private view of the cycle, because the cost of prudence is borne alone and the benefit is shared. The instrument that resolves this is not virtue but a common floor, imposed simultaneously and monitored. The letter asks the state to write an option that no private counterparty can write for itself, and for precisely the same reason.

The signatures are the substance. Dario Amodei signed, alongside four of his Anthropic co-founders; so did OpenAI’s chief scientist and its chief research officer, Shane Legg of DeepMind, Ilya Sutskever, and John Schulman. An independent tally of the roll at 1,134 names counted 533 from Anthropic and 330 from OpenAI, with roughly a quarter signing anonymously; the count stood at 1,324 this weekend. And then something genuinely without precedent: within a day, OpenAI and Anthropic endorsed the statement in their corporate names. The 2023 pause letter was signed by individuals, some of them executives acting personally. This one has been adopted by the institutions. An industry that has spent three years arguing it should be left alone has now, in its corporate voice, asked to be paced.

Cheaper, not smarter

Now the OpenAI announcement, where the popular account has slipped. The version circulating — that the flagship Sol improved its stablemates Terra and Luna — is wrong in a way that matters. What OpenAI’s engineering post actually describes is Sol being applied, “within a human-led process”, to the cost of running itself. Working through the company’s coding harness, the model rewrote the low-level kernels that execute its own mathematics, designed and ran hundreds of experiments on token generation, and monitored live training runs, intervening when they failed. The reported results: a 20 per cent reduction in the end-to-end cost of serving the model, and better than 15 per cent improvement in token-generation efficiency. The next day the savings were passed on — Luna down 80 per cent to twenty cents per million input tokens, Terra down 20 per cent, Sol’s own price untouched.

Follow the chain. The model improved its own serving stack, and the savings funded price cuts on the cheaper tiers. It did not make Terra and Luna better models, and nothing published describes any model choosing what its successor should be. This is a loop in economics, not in capability — real, compounding, and categorically different from the threshold the letter is aimed at. The figures, it should be said, are OpenAI’s own production measurements, audited by nobody outside the building.

For readers of Follow the Tokens, the repricing is the consequential item anyway. On OpenAI’s own benchmarking, its cheapest model now beats last year’s frontier at roughly six cents on the dollar per task; costs falling by that order change which business cases clear the hurdle without any change in what the technology can do. That is a repricing, not a revolution — and repricings are the thing most reliably mistaken for revolutions.

The doing

The stronger evidence sits in a paper Anthropic published on 4 June, which supplied the letter’s intellectual scaffolding and this essay’s title. Building a frontier model, it argues, divides into engineering — writing the code, standing up the infrastructure, overseeing the training — and research: deciding which experiments to run, which results to trust, which ideas are dead. Call these the doing and the deciding. On Anthropic’s own evidence, the doing is largely automated and the deciding is not.

The doing is where the numbers startle. As of May, more than 80 per cent of the code merged into Anthropic’s production codebase was written by Claude, against low single digits before Claude Code appeared in early 2025; the typical engineer now merges eight times as much code per day as in 2024. On the hardest, least-specified internal tasks, the model’s success rate reached 76 per cent in May — a rise of fifty percentage points in six months. On a standing test in which the model must make training code run faster while passing the same correctness checks, the best model went from roughly a threefold speedup in May 2025 to roughly fifty-two-fold this April; a skilled human, given an afternoon, manages about four. Externally, METR finds the length of task a model can complete unaided doubling every four months, up from every seven.

The paper marks its own numbers down as it goes, which is the reason to take it seriously. Lines of code measure quantity, not quality, so eight times overstates the true gain; the fifty-two-fold figure is flattered by how poor the starting code was, and the information is in the trend, not the multiple; a staff survey putting self-assessed uplift at fourfold is immediately discounted on the grounds that developers overestimate. A laboratory publishing competitive metrics in order to argue for its own restraint, and then deflating them, has earned a hearing. It has also met an old acquaintance: having multiplied the code flowing through the organisation, Anthropic reports that human review has become the new bottleneck, and that it now generates far more ideas than it has capacity to pursue. Amdahl’s law, translated from processors to institutions — accelerate one stage and the constraint relocates, invariably to the reviewing and deciding stage, which was already where the scarcest people sat. Any firm proposing to put agents inside a control environment should budget the review capacity before the generation capacity.

The deciding

And the deciding? Here Anthropic’s own experiments cut both ways, and the second is the one worth memorising. In April, Claude-powered agents were handed an open problem in AI safety and left alone with it. Over 800 cumulative hours and some $18,000 of compute, they recovered 97 per cent of the achievable gap; two human researchers, given a week, managed 23. Then the caveats, which carry the weight: the humans chose the problem and wrote the scoring rubric, and the result did not transfer cleanly to production scale. Within those bounds the agents designed every experiment themselves, and direction-setting was the only meaningful human role. Which is the point. It was still the human role.

The second experiment took 129 real research sessions at the exact moment a human researcher had taken a wrong turn, showed the model only the work up to that point, and asked what it would do next. In November the best model beat the human’s actual choice 51 per cent of the time; by April, 64. Read alone, that is judgement arriving. But the paper also ran the control, on 127 moments where the human’s next move was already strong — and there the model won only about 20 per cent of the time. That control is the most informative number in the whole exercise and the least reported. The models have become excellent at rescuing a session that has gone wrong, and remain some distance from beating a researcher who is already on the right track. Real progress; not yet research taste.

The letter and the prospectus

The statement deserves its sceptics, and they hold three good cards. Incumbents absorb compliance costs more easily than challengers, so a governance regime shaped by the two laboratories with the most to protect will tend to entrench them; the charge of capture is unproven but not unserious. Open-weight models — the insurgency examined in Kimi K3, Weighed — cannot be governed by any mechanism that reviews weights before release, and the sharpest competitive pressure on the American laboratories now comes from exactly that direction; a pacing mechanism that binds only the willing redistributes the frontier rather than pacing it. And the verification problem is stated most brutally by Anthropic itself: a training run is far easier to conceal than a missile silo, its inputs are general-purpose, and whoever defects quietly while others pause inherits the lead. The world has built verification regimes for dangerous technologies before. They took decades, and decades are what nobody involved believes are available.

One coda belongs on the record. Anthropic published its recursive self-improvement paper on 4 June — three days after confidentially filing a draft S-1 with the Securities and Exchange Commission, itself days after raising $65bn at a $965bn valuation. The same institution is arguing that its technology may need to be slowed while preparing to sell shares in the enterprise that builds it. The positions are not necessarily inconsistent. But a prospectus is a document in which risk factors must be disclosed, and this one promises to be an unusual read.

Keep the two claims apart, then. The doing is automated and accelerating; the deciding is not, and the distance between them is currently measured only by the laboratories themselves, on data only they can see. Read properly, the letter asks for one thing: that the instruments for measuring that distance, and the brake that would depend on them, be built by somebody other than the people being measured. On the evidence of the last week that is not alarmism. It is segregation of duties.