Ask a general reader who makes artificial intelligence and you will get two names, perhaps three, and a ranking: who is cleverest this month. The framing is understandable — a race is legible, and it has protagonists — but for anyone deciding what to actually build on, it obscures more than it reveals. The field is better understood not as a podium but as a market with segments: a dozen serious laboratories on three continents, optimising for materially different things. Knowing the players, and what each is actually playing for, is the first piece of literacy; seeing past the frontier to the quieter engineering beneath it is the second.

The visible frontier

The race the press does cover is real, and three laboratories trade its lead. Anthropic’s Claude, OpenAI’s GPT and Google’s Gemini have alternated at the top of the independent rankings for two years, and all three have refreshed their flagships within the past eight weeks: Anthropic’s Claude Fable 5, announced in June, introduced a new Mythos-class tier above its established Opus, Sonnet and Haiku lines; OpenAI’s GPT-5.6 arrived in July in three variants — Sol the flagship, Terra the intermediate, Luna the fast and cheap; and Google used its I/O conference to launch Gemini 3.5, pitched squarely at agentic work atop its Pro, Deep Think and Flash tiers. Behind them runs a well-funded fast-follower: xAI, whose Grok 4.5, released in July, was pitched by Elon Musk as comparable to Anthropic’s Opus-class models but faster and cheaper — and whichever way one weighs that claim, independent testing now places it among the top handful of models on composite intelligence indices. A firm that dismissed Grok in 2024 is entitled to a second look in 2026.

One marker of where the frontier has reached deserves note: it now arrives with government attention. OpenAI staggered GPT-5.6’s public release while US government partners assessed the models’ safety — capability in biology and cybersecurity being the stated concern — and Anthropic’s most capable tier, Mythos 5, is restricted to approved programmes rather than offered for general use. The frontier is no longer merely a leaderboard; it is becoming a regulated altitude.

The open-weight insurgency

Beneath the closed frontier runs the open-weight field — models whose parameters can be downloaded and self-hosted, the territory an earlier piece in this series mapped. Its geography has shifted in a way few predicted. Meta, which lit the open-weight fire, has gone quiet: its Llama 4 models of April 2025 remain the current generation as of mid-2026, the promised flagship Behemoth has never shipped, with reporting pointing to a delayed or paused launch over capability concerns, and the company’s frontier attention has reportedly moved to a closed line of models. The pioneer has paused while the field it created runs on.

Europe’s sole frontier laboratory, Mistral of Paris, is running a different race altogether. Its flagship Large 3, released in December 2025 under the permissive Apache 2.0 licence, is the largest open-weight mixture-of-experts model from a major lab, priced far beneath the American flagships — and if independent evaluations place it well behind the closed frontier on the hardest reasoning tests, that rather misses what it is for. For firms with European data-residency requirements, Mistral is frequently the default shortlist; sovereignty and cost are the product, and a larger successor entered early access in July on the back of a multi-billion-euro data-centre buildout.

It is the Chinese laboratories, though, that now set the open-weight pace, each with a distinct franchise. DeepSeek’s V4 line is the cheapest capable generalist; Alibaba’s Qwen is the most widely adopted base for building on; Zhipu’s GLM leads the Chinese field on coding; and Moonshot’s Kimi K2.6 is engineered for long agentic runs — the last of these currently the strongest open-weight model on composite indices, ahead of anything American labs have published openly. Baidu’s Ernie completes the set with a different headline entirely: Ernie 5.1, released in May, claims flagship-class performance at roughly six per cent of the pre-training cost of comparable models. Chinese laboratories now hold most of the top open-weight positions, with the best of them trailing the top proprietary models by single digits on composite benchmarks — a gap that has closed faster than most forecasts allowed.

Under the hood

All of which is still, in the end, one race — capability — measured two ways. The thesis of this piece is that it is not the only race being run, and for most enterprise purposes not the decisive one. Beneath the frontier, the industry’s engineering effort has swung toward a different set of axes: cost per token, latency, model size, specialisation, deployability, and the efficiency of training itself.

The first evidence is hiding in plain sight: every frontier laboratory runs an efficiency line beneath its flagship. Anthropic’s Haiku, Google’s Flash and its open-weight Gemma family, OpenAI’s Luna and mini tiers and its open gpt-oss models — these are not afterthoughts but volume products, the models that actually carry most production traffic, at a fraction of flagship prices. The flagship is the shop window; the small tiers are the shop.

The clearer evidence is two American giants entering the field this spring precisely there, rather than at the frontier. Nvidia — the company that sells the industry its hardware — now publishes serious models of its own: the Nemotron 3 family, in Nano, Super and Ultra sizes, open models aimed squarely at agentic systems, engineered above all for throughput — the Ultra model, announced at Computex in June, is the most capable American open-weight model and serves at several times the speed of comparable Chinese peers. The strategic logic is transparent and sound: models tuned to run agents cheaply and fast make the silicon beneath them more valuable. Microsoft, meanwhile, used its Build conference to unveil MAI, a family of seven in-house models each specialised for a single capability — reasoning, coding, image generation, transcription and voice — trained from scratch on its own silicon, and deliberately sized: its reasoning flagship competes, in Microsoft’s own framing, in the medium weight class, punching above it. Neither company is trying to out-frontier the frontier. Both are betting that the right-sized specialist, cheap to run and easy to place, is where the enterprise money is.

Run the eye back across the whole field and the same grammar recurs: sparse mixture-of-experts designs that activate a fraction of their parameters per token; hybrid architectures and aggressive low-precision formats built for inference speed; Baidu’s six-per-cent training claim; Mistral’s small reasoning models; task-specialists for code and voice. The industry’s unit of progress is quietly changing — from parameters to performance per pound, per millisecond, per watt.

Reading the map

For a finance function, the practical conclusions fall out directly. The first is that the market now has genuine segments, and the buyer’s question is which segment a workload belongs to — not which champion to back. This column argued last time that capability belongs to the whole assembly and that there is no model best at everything; the map above is that argument’s supply side. The frontier tier exists for the hardest problems: deep research, complex agentic work, the tasks where failure is expensive and capability is the binding constraint. The under-the-hood tier — the Haikus and Flashes, the Nemotrons and MAIs, the efficient Chinese open models — exists for the volume: classification, extraction, summarisation, the everyday work where cost per completed task, latency and governability decide. A firm that routes its volume through a frontier flagship is overpaying for headroom it does not use; one that gives its hardest problems to a budget tier has underbuilt.

The second conclusion is that optionality is real and increasing. On every axis — closed frontier, open weights, efficiency, specialisation, sovereignty — there are now multiple credible suppliers on more than one continent. For a sector whose supervisors have begun naming concentration in model supply as a systemic concern, that plurality is not a curiosity; it is a risk-management resource, there to be used.

The race the press reports is real, and it matters. But it is one race of several, and the quietest ones — cheaper, smaller, faster, closer to the metal — are where most of the field’s engineering now lives, and where most enterprise value will be captured. The buyer’s edge is not picking the winner of the race everyone is watching. It is knowing which race their workload is actually running — and noticing how much of the field is under the hood.