Copilot can't read a national annex

It can't, and neither can ChatGPT or Claude out of the box. Copilot in the title stands in for a category: general-purpose AI dropped into regulated European engineering. Five properties of the documents defeat that whole category, one by one.

An engineer reading a drawing at a desk by a window overlooking the city

“Like an untrained intern”

The sharpest verdict on generic AI in engineering comes from the people who deployed it. At a large engineering consultancy, the person who led the Copilot rollout summed up a year of company-wide use in one line:

Hopplöst för allt specifikt — den vet inte hur vi jobbar. Som en otränad praktikant.

The person responsible for the Copilot rollout at a large engineering consultancy. “Hopeless for anything specific — it doesn't know how we work. Like an untrained intern.”

That is not a Microsoft problem. Copilot runs on the same class of models as ChatGPT and Claude, and every general-purpose tool built on that class fails the same way here. A bid team at another firm said the same of ChatGPT: the answers are always mixed; you always have to double-check. The models are trained on the public internet, and the documents that govern a European infrastructure project are the part of the written world the public internet covers worst. The failures are specific and repeatable. Here are the five that matter.

The annex that silently overrides the code

The Eurocodes look like the perfect AI use case: harmonised European standards, widely discussed, well represented in training data. The trap is that they are deliberately incomplete. Each country sets its own values for the nationally determined parameters: partial factors, snow and wind loads, exposure requirements. Those choices live in a national annex, and the annex quietly overrides the base text.

Ask a general model a question under EN 1992, the concrete standard, and it answers from the base document: fluent, plausible, and calibrated to the recommended values rather than the Swedish ones. Nothing in the base text tells the model an override exists. For engineering, that is the worst failure class there is: not a refusal, not nonsense, but a confidently wrong parameter that looks exactly like a right one.

One project, two languages

A Nordic project set rarely stays in one language. The contract is in Swedish, a technical specification in English, environmental permit conditions in Swedish, a supplier's design basis in English. The requirements are one body of requirements; the words are split across two languages.

Generic retrieval treats language as a relevance signal. Ask in English and the English documents come back; the Swedish clause that contradicts them stays buried. Cross-references break the same way: a Swedish document pointing at section 4.2 of an English one is a link no chatbot follows. The answer is not wrong because the model cannot read Swedish. It is wrong because the model never realised the Swedish half was the half that mattered.

Rules that only exist as scans

Infrastructure has a long memory. A structure built in 1968 is assessed against documents from 1968: older design standards, water rights judgments, local regulations, archived permits. Much of that still governs projects today, and it exists only as scanned paper: typewritten pages, stamps, tables, notes in the margin.

General tools run cheap OCR and hope. On clean office documents that works; on a scanned table of load values it drops superscripts, merges columns, and turns limits into different numbers. The model then reasons fluently on top of the corrupted text, with no idea it is reading garbage. The scan does not need to be exotic to cause this. One bad table is enough.

Words the training data barely contains

Swedish infrastructure runs on a compact technical vocabulary with contractual weight: bygghandling, systemhandling, förfrågningsunderlag, granskningsutlåtande, ÄTA. These are not words waiting to be translated; they encode status. A systemhandling and a bygghandling can describe the same object, and the difference decides whether a requirement is binding yet.

There is very little public Swedish engineering text for a foundation model to learn this from. So a general model translates instead of understands: bygghandling becomes "construction document", the status distinction evaporates, and an answer that reads smoothly in English has quietly flattened the one distinction that made the question worth asking.

The same code, a different meaning

AMA is the reference framework Swedish technical descriptions are written against, and it is published in editions: AMA 17, AMA 20 and onward. A code invoked in a technical description means what the cited edition says it means. Between editions, requirements get revised, moved, and re-scoped, while the code itself stays the same string of characters.

A general model has seen fragments of AMA discussion online, from mixed editions, with no concept of which edition governs. Ask what a code requires and you get a blend: mostly one edition, seasoned with another, served as a single answer. The governing document declares its edition up front. The only correct behaviour is to verify every answer against that declared edition, and a chatbot has no mechanism for doing so.

One root cause, five symptoms

All five failures share a root. The right answer depends on which document governs: which annex, which language, which scan, which term, which edition. A general-purpose model has no way of knowing that, and worse, no habit of checking. Better prompts do not fix it. The intern is not lazy; it is untrained, and the training it lacks is not on the internet.

That is why this is a category difference, not a product gap the next Copilot release closes. A tool that works in this domain has to know which documents govern a project and verify what it says against them before it speaks. What that difference looks like in practice is the subject of a companion piece: a chatbot answers a question, Yesper does the work

Benjamin Glaser Co-founder at Yesper. Writes about AI and the industry that builds the world. benjamin@yesper.ai

Yesper is the AI civil engineer for construction and infrastructure. AFRY, COWI, NRC Group and other Nordic firms use it to halve the time on a study, rerun calculations in minutes, and catch errors that would otherwise slip through. Get in touch if you'd like to see what it can do for you.

Book demo