In the last week of January 2025, a laboratory in Hangzhou released a reasoning model that performed within reach of the best American systems and appeared to have been trained at a fraction of their cost. The reaction was not admiration but alarm. Within days DeepSeek had topped the app charts on both sides of the Atlantic; within weeks, governments began to act. Italy blocked the service; Australia barred it from federal devices; Taiwan, South Korea and India followed. In the United States, agencies from the Department of Commerce to NASA removed it from official systems, more than a dozen states did the same, and Congress received a bill with the unimprovable title of the No DeepSeek on Government Devices Act.

The worry deserves more precision than the coverage gave it. Four distinct claims are habitually run together: jurisdictional — data sent to the service is stored in China, where the state may compel its surrender; architectural — its data paths ran to other Chinese firms; operational — an exposed database, guardrails easily circumvented; and, quieter but more enduring, a question of values — like its domestic peers, the model deflects on subjects the Chinese state finds sensitive. Notice what the first three share. Each is a claim about a hosted service — an application that carries a user’s words across a border. None is a claim about the model. What was banned in Rome, Canberra and Washington is an application and its data path, not the mathematical object beneath it. The two are separable, and that is the whole point.

The farm and the kitchen

Before the vocabulary, a picture — aptly chosen, because “farm to fork” is the language of food traceability, and traceability is precisely the axis on which this debate turns. The training data are the seeds and the soil. The training run is the growing season: the costly, months-long transformation of raw inputs into something harvestable. The weights — the billions of learned numbers a trained model comes down to — are the harvested crop: durable, transportable, inert on its own, but the thing of value. The architecture is the recipe card; the runtime is the kitchen, commonplace and free; and the meal is inference, the answer one actually consumes.

An open-weight release hands you the harvested ingredients and the recipe card — and kitchens are everywhere — so you can cook the dish and season it to your taste. What you are not given is the farm: not the seeds, not the field, not the record of how the crop was raised. Here the Chinese question resolves into two things that must be kept apart. You still cook in your own kitchen: the weights run on your premises, nothing you plate is sent back to the farm, and so the data-exfiltration worry — the one the bans are about — does not survive self-hosting. But the character of the crop was fixed in a growing season you never saw. The model’s silences and slants were impressed on its parameters, and they travel with the weights into your kitchen. The provenance gap is real; it is simply not the data gap, and conflating the two is the central error of the public debate.

Open weights, not open source

That the Chinese frontier is released this way — open-weight, under permissive licences such as MIT and Apache 2.0 — is less a cultural preference than a response to constraint. Denied the most advanced American processors by export controls since October 2022, Chinese laboratories were pushed towards efficiency; openness then compounds the strategy, building an ecosystem and goodwill while eroding the position of Western labs that keep their models closed. Inference on the leading open Chinese models can run at a fifth, sometimes a thirtieth, of the cost of the closed Western flagships. And because the models are open-weight, an institution need send nothing to Hangzhou to use them: it downloads the weights and runs them inside its own perimeter. Once downloaded, they are permanent — no developer in China can revoke, meter or switch them off. The umbilical cord to the foreign service, the very thing the bans concern, has been cut.

The distinction the trade press blurs is this. Releasing a model open-weight publishes the learned parameters — closer to a compiled binary than to source code — while withholding the training data and the code that produced them. “Open source” is a stronger claim: the Open Source Initiative’s AI definition, published in October 2024, requires the freedoms to use, study, modify and redistribute, plus disclosure of the code, the weights and enough information about the training data to recreate an equivalent system — the farming diary, in the terms of the analogy. Most of what is marketed as open-source AI, Meta’s Llama as much as the Chinese frontier, is on inspection only open-weight. The gap is one of verifiability, not usability. Running a model needs only the small architecture file shipped with the weights and a freely available runtime — not the training code, which at millions of pounds a run almost no recipient could rerun anyway. The difference, in the end, is between what you can do with a model and what you can verify about it.

What a regulated firm should make of it

“Should we use Chinese models?” is the wrong question, because it collapses three matters a finance function is well practised at keeping apart: where a thing is deployed, where it came from, and what one is permitted to do with it. Disentangled, the decision becomes tractable.

On deployment, the gravest concern dissolves under self-hosting. For a firm bound by data-residency and operational-resilience obligations — the Prudential Regulation Authority’s expectations, the EU’s Digital Operational Resilience Act, domestic data-protection law — an open-weight model run inside the perimeter is, on this axis, more tractable than a closed model reached only through a foreign interface. The data never leaves; and because downloaded weights cannot be revoked, vendor-withdrawal risk does not arise.

On provenance, self-hosting cures nothing. Model-risk discipline does not care where a model was made. The PRA’s supervisory statement SS1/23, which expressly reaches artificial-intelligence and machine-learning models, and the SR 11-7 canon that governs the same territory in the United States, require that any model bearing on a material decision be inventoried, validated, monitored and owned — and opaque provenance makes each of these harder.

Who, inside a firm, does this fall to? Not the team that deploys the model but the second line of defence: independent model-risk management, whose job is to challenge what the first line builds, with internal audit assuring the framework as a whole. Validating a model of undisclosed provenance, though, inverts the usual method. The build cannot be reproduced, because the data and training code are withheld; validation becomes empirical rather than constructive — characterising what the model does rather than certifying how it was made. In practice that means pre-defined scenario suites drawn from the intended use; deliberate probing of the known failure modes — embedded content restrictions, bias, susceptibility to prompt injection; adversarial red-teaming; and benchmarking against a trusted reference or human judgement. The format confers one advantage a closed service never grants: with the weights in hand, the validator can inspect the model’s internals, pin the exact version under test, and rerun any check at will — no vendor can silently update the artefact beneath the audit.

The test should also fit the use — the proportionality model-risk frameworks already demand. A customer-portal voice assistant and an investment-screening model do not warrant the same scrutiny. For the assistant, the exposure is in what the model says: conversational red-teaming, confirming that political slants and refusals do not surface in the channel, checking guardrails, testing for hallucination and the leakage of sensitive data into a reply — a conduct and Consumer Duty lens. For the investment or credit screen, the exposure is in the decisions the model shapes: bias and fairness across customer groups, stability under stress, back-testing against realised outcomes and, above all, explainability — whether a model whose origins cannot be inspected can discharge the duty to explain a material decision at all. The embedded restrictions that are a shrug in a coding assistant, and a nuisance in a customer channel, may be disqualifying in a lending decision. The machinery is the same; its depth is calibrated to materiality.

Set against the provenance burden is a genuine benefit: by letting a firm run a model itself, on infrastructure of its choosing, open weights reduce reliance on any single provider — real diversification at a moment when supervisors are naming concentration in model and infrastructure supply as a systemic concern.

The measured conclusion is that open weights are a real instrument for regulated finance, precisely because of the distinction drawn here: they let an institution separate the useful artefact from its untrusted origin — take the crop into its own kitchen while declining to trust the farm. But the same distinction forbids the opposite error. Open is not audited. The weights carry their farm with them, in behaviours no one specified to you and no one can fully show you; a model of undisclosed provenance that bears on a regulated outcome must be validated, contained and governed with that uncertainty in full view.

There is an older instinct in this profession that fits the moment exactly. Long before there was a model to weigh, finance turned on weights and measures — verifying the quantity, certifying the standard, trusting no consignment further than it could be tested. The open-weight era does not retire that instinct. It restores it. A downloaded model is a consignment of uncertain origin: immensely useful, freely portable, and to be trusted precisely as far as it can be measured — and no further.