Some years ago, when my eldest was three or four, I telephoned home from the office one afternoon. He had not long been back from nursery, and I asked him — in the way one asks a small child — what he was doing. I was expecting an answer somewhere between an episode of Thomas the Tank Engine and a plate of fish fingers. What I got, after a short pause, was a deadpan two words: “Talking to you.”

He was, of course, entirely right. He had answered the precise question I had asked — what are you doing? — with the most immediate and truthful answer available to him. The fault, such as it was, was mine. The question I uttered was not the question I intended. I meant something closer to “what were you up to this afternoon, before I rang” — a whole freight of context I was carrying in my head and had not troubled to put into words. He had no way to reach it, so he answered what was in front of him.

A three-year-old grows out of this. Within a year or two he would have inferred what I meant — read the situation, filled in the unspoken, supplied the context I had left out — because that is what human intelligence learns to do: it reads the room. A large language model does not grow out of it. It answers the question you actually asked, with exactly the context you actually supplied, and not one word more. The literal-mindedness is not a phase to be outgrown with a few more months of training; it is a permanent property of the thing. And the whole discipline of using these models well is, at bottom, the discipline of engineering around it — along two axes that those two words on the telephone happen to capture exactly.

Say what you mean

The first is the one my son illustrated directly: the model does what you say, not what you mean, so you had better say what you mean. This is what the faintly grand term “prompt engineering” describes, and it is far more intuitive than the name suggests.

Asked to “summarise this report”, a model will certainly produce a summary — but at what length, for whom, foregrounding what? Those choices do not vanish when you decline to make them; they are simply made for you, and rarely as you would have made them. Ask instead for “a summary in fifteen lines for a risk committee, leading with the findings that bear on capital”, and you get something you can use, because you have said what you meant. It also explains a puzzle every organisation eventually meets: why the same model, on the same day, delights one colleague and disappoints another. The difference is frequently not the model at all, but the asking — one has learned to state what she means; the other issues a terse instruction and faults the tool for reading it as written. The craft is not a book of incantations. It is the plain, unglamorous habit of making the implicit explicit: naming the role you want the model to take, the audience it is writing for, the shape of the answer you need, the boundaries it must not cross, and — where it earns its keep — a short example of a good answer, so the model can see the target rather than guess at it. None of this is clever. It is merely the refusal to assume a machine can read a mind you have not written down.

Put the room in front of it

But return to the telephone, because the anecdote holds a second lesson beneath the first. Why could my son offer only “talking to you”? Partly the literalism — but partly because a telephone call is a narrow pipe. I could not see the room: the abandoned toys, the half-watched television, the debris of lunch. He answered from the thin slice of the world the channel let through. Had I been standing in the doorway, I need not have asked at all; the context would have been in front of me.

A model labours under precisely this constraint, and it is the second great lever. Whatever it absorbed in training, in the moment it can reason only over what is placed before it — its context window, the span of text it can hold in view at once. What falls outside that window may as well not exist. Give a model a question with no background and it answers, like a child on the telephone, from an impoverished view: thinly, literally, technically correct and practically useless. Give it the relevant policy, the actual numbers, the prior correspondence, the definitions your own firm uses, and the same model reasons richly — because you have put the room in front of it.

There is a subtlety here that the analogy sharpens rather than hides. More context is not simply better. Had my son recited the entire day — every toy, every snack, the plot of every cartoon — I should have struggled to find the answer I wanted in the torrent. A model is no different: bury the one relevant figure in a hundred pages of irrelevance and its answer degrades, because it must now find the signal you failed to isolate. The aim is not the most context but the right context — a discipline of retrieval and curation, not of volume. This is why context length has become one of the specifications the laboratories now compete on, and why the real labour of deploying these systems has less to do with the model than with what surrounds it: assembling the right documents, retrieving the passages that matter, holding the thread of a long task so the beginning is not lost by the end. A context window, like a telephone line, has an edge — beyond which the earliest things said begin to fall away, and past which even the ablest model is once again guessing in the dark.

Fit for a decision

For a finance function, none of this is developer fiddliness to be nodded through as somebody else’s affair. It is the difference between an answer that can be relied upon and one that is confidently, precisely wrong in a way that reaches a decision. This is the particular danger of a literal machine: asked a carelessly framed question, or starved of the context a judgement requires, it does not fail loudly or leave a blank. It fails in the most treacherous manner available — returning a fluent, plausible, immaculately formatted answer to a question subtly different from the one intended, or reasoned over a fraction of the relevant facts. The failure wears the costume of success.

The earlier pieces in this series held that a model bearing on a material decision must be validated, monitored and owned, and that capability is a property of the whole assembly rather than of the weights alone. Prompt design and context are where that principle meets the desk. How a question is put to a model, and what the model is given to answer it with, are as material to the outcome as the choice of model itself — and, in a regulated setting, as deserving of standardisation, review and control. There is a further reason they belong in scope: how an answer was reached is part of the answer. The question put and the material placed before the model are the inputs to a decision, and inputs to decisions are meant to be recorded, versioned and reviewable — not improvised at a keyboard and forgotten the moment the reply appears. A prompt is not a private keystroke; where it bears on a regulated outcome, it is part of the audit trail.

My son, I am glad to report, learned long ago to read a room, and would now find the question faintly absurd. The models will not learn it. They remain, and will remain, superb literalists — answering exactly what is asked, from exactly what is before them, with a three-year-old’s precision and none of the growing up. The skill is not in wishing them otherwise. It is in learning to ask the question one actually means, and to place within their view the world they need in order to answer it — so that “talking to you” becomes the opening of a useful conversation, rather than the whole of it.