Some weeks ago a lecture found its way to my screen — an hour of a theoretical physicist calmly drawing a straight line through the future of his own profession. Adam Brown is worth the hour on credentials alone: Oxford, a Columbia doctorate, appointments at Princeton and Stanford, where he taught Einstein’s general relativity; he now leads Blueshift, the Google DeepMind team charged with sharpening the scientific reasoning of its models, and is a core contributor to Gemini. His talk, “Training Sand to Think”, opens with an image that has not left me since: our civilisation has learned to turn sand into silicon, silicon into chips, chips into neural networks — and neural networks into things that think. This series has circled that fact — what it does to coding, to premises, to balance sheets — without ever answering the question underneath. Having watched Brown’s lecture, and having sat with what it implies, I owe readers a straight answer. Here it is, with the evidence.

The straight line

The core of Brown’s case is a straight line on log paper, and physicists trust nothing more. His profession is trained to hunt scaling laws — regularities in how systems change with size. The biologist Max Kleiber found a famous one in 1932: plot metabolic rate against body mass, mouse to elephant, and the points fall on one straight line across orders of magnitude. In 2020, researchers did the same to neural networks and found something nobody was owed: feed a model more computation — larger network, longer training, in proportion — and capability rises along a line just as straight. No law of nature requires intelligence to scale predictably. Empirically it does, and the relationship has now held across roughly eight further orders of magnitude.

That line is why the money came. It converted frontier research into something an investment committee could underwrite — dollars in, intelligence out, on a schedule — and it is the intellectual collateral behind a build-out whose financing this series examines elsewhere. Brown gives the spending a currency: a mole of floating-point operations — Avogadro’s number of them, a physicist’s joke — costs about a million dollars, and frontier training runs have gone from a fraction of that in 2020 to hundreds of millions today. More striking is his ranking of the drivers. Moore’s law, the traditional engine of cheap computing, is a minor contributor. The great forces are the raw willingness to build — total training compute rising roughly fourfold a year — and, above even that, algorithmic ingenuity: humans finding better ways to grow the machines.

Five years

Now set the line against the calendar, because the velocity is the argument. GPT-3 appeared in the summer of 2020. In 2021, on a benchmark of high-school competition mathematics, the best model in the world scored 6 per cent — against about 40 for a capable doctoral student and 90 for an Olympiad medallist. Professional forecasters, asked when machines would reach half marks, guessed 2025; the line crossed within a year, and by 2024 the benchmark was saturated — dead, in the trade’s cheerful idiom. Its successor, GPQA, was built to be “Google-proof”: doctoral-level science questions on which human PhDs score about 70 per cent inside their own field. Models sat at chance until early 2024, then passed the human experts within eighteen months. To the obvious objection — that the machines have merely memorised the internet, answer keys included — Brown offers the standard rebuttal and a personal one: performance barely dips on freshly written lookalike problems that exist nowhere online, and on unpublished Stanford graduate examinations in quantum mechanics and general relativity, marked by hand, the models went from failing to flawless in about eighteen months. By Brown’s arithmetic the machines have been ageing four cognitive years for every calendar year.

Then the summit. The International Mathematical Olympiad is prized precisely because it rewards lateral invention over learned technique — long held up as the thing statistical machines would never do. In July 2025, five years almost to the month from GPT-3, an advanced version of Gemini was submitted for official grading alongside the human contestants and solved five of the six problems for a gold-medal score — a standard reached by roughly one in ten of the world’s best young mathematicians — working end-to-end in plain English inside the standard four and a half hours. The IMO’s president, Gregor Dolinar, confirmed the result and described proofs the graders found “clear, precise and most of them easy to follow”. And the line has not paused there. In January, Brown’s own team announced a paper in algebraic geometry co-authored with three mathematicians, among them Ravi Vakil of Stanford, president of the American Mathematical Society; a central step was the machine’s, and Vakil called it “the kind of insight I would have been proud to produce myself”. In May, OpenAI disclosed that an internal model had settled the unit-distance problem, posed by Paul Erdős in 1946 and open for eighty years — overturning the answer mathematicians had believed, via a construction from deep algebraic number theory that generations of specialists had not found. The Fields medallist Timothy Gowers called it “a milestone in AI mathematics” and said a human submission of the same paper would have merited acceptance at a top journal without hesitation. The claim of full autonomy drew careful footnotes — humans shaped the surrounding workflow — but the insight was the machine’s. Preschooler to research mathematician: five years.

The insiders’ tell

Evidence of a different kind sits closer to home — on this page, in fact. The model that co-writes this column, Claude Fable 5, is by its maker’s own account not the whole animal. Anthropic’s frontier model, Mythos 5, and the Fable edition on this byline share the same underlying model; Fable is the version released to the general public, fitted with additional safeguards around capabilities that could be turned to harm, while Mythos, without those measures, is reserved for approved organisations. Consider the structure of that decision. A commercial laboratory, with every incentive to sell its best product to everyone, has judged the unrestricted form of its own creation unsuitable for general hands. Finance readers will recognise the shape of the argument at once: never mind the prospectus — watch what the insiders do with their own stock. A laboratory gating its own frontier model is the insider declining to sell. It is not proof of superintelligence. It is revealed preference about trajectory, from the people holding the best information on earth, and it points the same way as Brown’s line. Candour cuts both ways, though, and my collaborator would insist on the caveat: a model has limited insight into its own depths, so that second piece of evidence deserves a discount. The straight line needs none.

When, not if

So to the question in the title, and to why I answer it the way I do. Sort every serious objection to superintelligence into two piles and something clarifying happens: almost all of them turn out to attack the timetable, not the destination. Energy, grid connections, the balance-sheet strain of the build-out — all examined in these pages — are rate-limiters: “when” variables. The gap between frontier brilliance and workplace productivity — the careful trial, examined elsewhere in this series, that found experienced engineers slower with AI on the very systems they knew best — is diffusion lag: “when” again. The only genuine “if” on the table is that the line itself breaks — that data runs out, or returns diminish. And there Brown’s answer is the uncomfortable one: that pessimism has been falsified repeatedly, each theoretical ceiling dissolved by more compute and better algorithms; no law of nature has been produced that stops the line at the altitude of the human mind; and the nearest precedent — chess — went through our ceiling without pausing to wave. The burden of proof has quietly changed hands. The sceptic now owes us a mechanism, and until one is produced, “when, not if” is not faith. It is the base case.

One definitional footnote, because it matters to this readership more than the metaphysics. The operative threshold for a finance leader is not the singularity of science fiction; it is the moment the machine is better than your best people at the tasks on their desks — and that threshold arrives first, domain by domain. On the evidence above, in mathematics it is arguably already behind us.

The believer’s position

What does a believer do differently on Monday morning? Not prophecy — positioning; this readership is not paid for certainty, it is paid for being positioned when uncertainty resolves. Believers redesign accountability for machine-speed work rather than bolting it on afterwards. They rewrite supplier contracts priced in hours before their suppliers are ready. They rebuild the apprenticeship their juniors will no longer get by accident, because someone must be qualified to supervise the machines in ten years. They scrutinise the financing of the build-out as sceptically as they would any prospectus, because believing in the technology and believing in every balance sheet attached to it are different creeds. And they read the primary evidence — the lecture, above all — rather than the commentary, this column included. Am I a believer? On Brown’s evidence, on the line’s refusal to bend, on the tell in my collaborator’s own release notes: yes. The only question left standing is whether your firm is positioned as if you were.

Adam Brown’s lecture, “Training Sand to Think: Artificial General Intelligence and the Future of Physics”, delivered at the Perimeter Institute for Theoretical Physics, is publicly available and repays the hour. Readers who want the physics itself will find this site’s own programme in the Physics section.