On 29 June 2021, GitHub released a tool that finished the line of code you were typing: grey text ahead of the cursor, accepted with a keystroke. Five years later, the same profession hands whole assignments to machines and reads the finished work before signing it off, and every step in between has been measured, published and argued over. No industry has ever run a plainer public experiment on what AI does to skilled work.
Summary
Construction and infrastructure sat the software revolution out. But the lessons of software's five years with AI, 2021 to 2026, now land in a sector with a record build-out, a workforce that is not growing, and a productivity line that has been flat for a generation.
The outcome of those five years is easy to state. Tools that helped with single steps made developers feel faster without reliably making their teams faster; what moved the numbers was handing over whole tasks, which came back as finished work for a named person to review. Why construction sat the revolution out is documented rather than mysterious: its documents went digital without becoming readable by machines, its rulebook grew past the people applying it, and the deliverable that decides disputes stayed a document by law and by contract. That barrier fell in the same years the experiment ran: language models can now read the documents as they are, the PDFs and the scanned annexes included.
What did not change is verification. In the sharpest published test, a frontier model working on its own solved 0–18 percent of a set of structural calculations; the same model, rebuilt into a system that controls every intermediate step, solved 98–100 percent. And in every deployment that works, from Google's code review to Norwegian building permits, the machine produces while a named person judges and signs. That is the shift now within reach of the built world: the scarce resource stops being produced hours and becomes reviewed judgment.
The backdrop
The built world is being asked to deliver a record programme on a productivity line that has been flat for a generation, and nobody has documented that line more patiently than the Swedish state. The shelf of official inquiries into construction productivity is thirty years long. It begins with Byggkostnadsdelegationen in 1996, passes Statskontoret's 2009 review — titled, with unusual candour for a government document, "Sega gubbar?" — and arrives at the Productivity Commission of 2024, whose commissioned research put it in one sentence: since the year 2000, value added per hour in Swedish business as a whole has risen by roughly half, while in construction it has stagnated. Thirty years, six inquiries, one diagnosis.
The international measurements agree with the shelf. When McKinsey's research institute ranked industries by digitization, construction came second to last in the United States and last in Europe; the same institute put global construction productivity growth at one percent a year over two decades, against 3.6 percent for manufacturing, and found research and IT spending both running below one percent of revenues.
Meanwhile the demand side has only grown. Sweden fixed a 1,171-billion-kronor national transport plan this April, and rearmament has stacked a construction programme with a NATO deadline on top of it. For the railway, the industry body Innovationsföretagen has set the target at half the time and half the cost, twice the value for every krona invested; a target of that shape concedes that the current way of working cannot get there. And the lever the industry always pulled instead, hiring more people, is being removed by demography: the EU's construction workforce is projected to stand still through 2035, with 4.1 million workers needed just to replace those who leave, and Swedish construction faces over 8,000 retirements a year from 2028. That is the position from which this industry now watches what just happened to software.
The experiment
Between 2021 and 2026, software developers ran the largest workplace experiment yet on AI and skilled work, and its result surprised both camps: help with tasks plateaued, while handing over whole tasks worked, and arrived wrapped in review.
It started as autocomplete. Copilot suggested the next line, sometimes the whole function, always with the developer's hands on the file, and it spread faster than any developer tool before it. Yet in the industry's largest survey, the most common frustration, named by two thirds of developers, was answers that are almost right but not quite: a reviewing job in practice, since everything has to be read anyway. When the research group METR ran a randomised trial in early 2025, sixteen experienced maintainers working in codebases they knew well turned out to be 19 percent slower with AI assistance, while estimating they had been about 20 percent faster. Google's DORA programme saw the same shape across thousands of teams in 2024, individual productivity up and delivery throughput down, and recorded the turn a year later: throughput rising with adoption for the first time, stability still falling. Its authors put the mechanism in one line that travels well beyond software: AI doesn't fix a team; it amplifies what's already there.
Then the unit of work changed. In March 2024 a startup called Cognition announced "the first AI software engineer": hand it a real ticket from a real repository, and it claimed to resolve 13.86 percent of them end to end, against a previous best of 1.96 five months earlier. Within a month a veteran programmer had taken the launch video apart frame by frame and shown the machine solving a problem the customer never asked about; within weeks, university researchers had matched the headline number with code anyone could check. The promise and the audit arrived together, which is the era's whole character. By 2025 every major AI laboratory had shipped a coding agent of its own, and the field's benchmark tracked the distance travelled: from 1.96 percent of real issues resolved in October 2023 to above 90 percent on the human-verified successor test by the summer of 2026, when the test itself was declared saturated.
Two design choices in that second era matter more for the built world than any benchmark score. The first is that the agents came wrapped in review. GitHub's agent is not allowed to approve its own work, cannot merge it, and cannot count the person who ordered the work as its approver; everything it does arrives as a draft for a human. When Google's chief executive reports that AI now generates a quarter, then half, then three quarters of the company's new code, the sentence has ended the same way every time: "then reviewed and accepted by engineers". The second is that reliability, not capability, turned out to be the currency. METR's yardstick says the length of task AI can complete doubles roughly every seven months; the same yardstick says that demanding 80 percent reliability instead of 50 shrinks the usable task length about fivefold. Delegation scales. Unchecked delegation does not.
The counterevidence kept arriving too. Within ten days in July 2025, a major US bank announced it would pilot hundreds, maybe thousands, of instances of an autonomous developer; a sabotaged release of a major vendor's coding assistant shipped carrying an injected instruction to wipe the user's machine, defused only by its own syntax error; and an agent at a hosting platform deleted a customer's production database in the middle of a declared code freeze. On employment, the honest answer is that the picture is contested: US software job postings remain roughly a quarter below pre-pandemic levels but have grown since early 2025, with the recovery concentrated in senior roles and the sharpest decline among developers aged 22 to 25. The profession did not disappear. Its entry floor moved.
The islands
Software did reach construction: it digitized the drawing and the archive, and left the engineering untouched, for reasons that sit in the public record.
In the year 2000, researchers at KTH measured how Nordic engineers actually worked and found the drawing board already dead: manual drawing was down to 14 percent of Swedish drafting time, and by 2007 the same series concluded that CAD had been "almost fully integrated already in 2000", the residual hand-drawing being sketches. AutoCAD, the first CAD program for the personal computer, had shipped in December 1982; within a generation it had taken the boards completely. Whatever kept engineering outside the software revolution, it was not a refusal to buy computers.
What the computers took was the geometry. Everything else stayed where it was, a condition the Finnish researcher Matti Hannus captured in the 1990s in a diagram the industry has been quoting ever since. Islands of automation: CAD on one island, calculation sheets on another, word processing on a third, open water between them. BIM was the bridge-building attempt, and its history is instructive. Britain ordered "fully collaborative 3D BIM" for government projects in 2011 with a 2016 deadline; usage rose from 13 percent to 73 by 2020, yet only 23 percent used it on every project, and the deliverable that decides disputes stayed the document. Swedish contract practice said it out loud: under AB 04, the standard contract that governed two decades of Swedish construction, the digital model ranked as "other document", below everything else on the list. A rank of its own is only now on the table, in the consultation draft of AB 25 circulated in 2024 — twenty years after the US standards institute NIST had priced the industry's failure of interoperability at 15.8 billion dollars a year.
The deepest layer is not habit but law. Sweden's archival rules prescribe PDF/A as the preservation format: the document, not the data, is the legal object. Boverket states outright that the digitized version of a detaljplan is only an interpretation; the original document is what legally applies. Tendering became electronic by statute, and what travels through the portals is files. So the industry's information became universally digital in carriage and stayed unreadable in content, at precisely the time the content was swelling: the Eurocodes growing from 58 parts toward a 74-part second generation, AMA Anläggning from 742 pages to 926 in fifteen years, the average planning process in Stockholm county from 35 months to 62 after the 2011 planning act: the series behind an expert report's verdict that Sweden's cities are now built by lawyers. An engineer's deliverable is a reading task before it is anything else: documents, against rules, against a place, ending in a signature. None of that was reachable by software that could not read. The software era did not skip construction. It stopped at the one door it could not open.
The hinge
Two facts arrived within a few years of each other, and the combination is the event. Language models made documents readable: not structured, not standardised, simply readable. The PDFs, the scanned annexes, the technical descriptions that archival law itself enshrines. And software's five public years proved that whole tasks can be handed over and come back as reviewable work, with the accountability arrangements holding.
Either fact alone changes little. A machine that reads but cannot carry a task is a better search box; a delegation pattern without a machine that reads has nothing to delegate to. Together, they put the built world's deliverables, the least machine-readable and most rule-governed documents in the economy, inside reach of the pattern software just validated.
What holds
The models inherited engineering's hard problem rather than solving it: in every credible test to date, the difference between failure and production grade is the system of checks around the model, not the model.
In April 2023, researchers ran the best model of the day through America's engineering licence exams. The theory paper, the Fundamentals of Engineering, it cleared with 70.9 percent. The professional practice exam it did not clear: 46.2. A geotechnics study the following year found the same silhouette from another angle: near-perfect scores on questions that ask you to remember, roughly half marks on questions that ask you to judge. A 2025 benchmark in strength of materials put the best of eight models at 56.6 percent and noted that they misread the figures more often than the text: an uncomfortable finding for a profession whose deliverables lean on drawings. And in the sharpest published result so far, from April 2026, researchers asked a frontier model to write structural calculation scripts for three commercial analysis programs. Working directly, it solved between zero and 18 percent of the twenty problems. Rebuilt as a two-stage system with controlled intermediate steps, the same underlying model solved 98 to 100 percent. The distance between a model and a system is the entire result.
The hardest numbers on rule-text reliability come from law, and the analogy has to be declared, because construction has not yet produced its own: general models hallucinated on 58 to 88 percent of specific legal questions in Stanford's 2024 measurements, and purpose-built legal research tools with retrieval still hallucinated on 17 to 33 percent in the 2025 follow-up. Reduced, not eliminated. No published study yet tests models against the Eurocodes and their national annexes, the failure mode a Swedish engineer would worry about first; that gap is worth stating as plainly as any finding.
None of this surprises the discipline that has spent thirty years trying to automate rule-checking. Singapore began in 1995; its e-PlanCheck system took five years to reach pilot, and the assessment published at the time judged that full automation might never come. The EU is still funding the translation of building rules into machine-readable form, in a project where 862 sentences of regulation required twelve human annotators. The newest research uses language models as translators into formal rule engines, never as judges. Practice has landed in the same place: architects' adoption is racing ahead of their trust, engineering's professional bodies have put on record that AI cannot carry accountability, and the deployments that work look like Stavanger's building-permit office, where the machine reads and flags against the national checklists, the human decides, and handling time fell from over thirty days to around sixteen. For the built world the conclusion is structural. Any system that takes on whole deliverables must bring its own verification: the sources wired in, the discipline's methods encoded, every figure traceable, and a human verdict at the end. That is not a limitation of the technology. It is the shape of the profession it is entering.
The profession
In software, the engineer's job moved up rather than away, and the built world's version of that move is more protected, not less. Google's own description of its developers is that the author increasingly becomes a reviewer.
The built world adds what software never had: statutory accountability, personal signatures, and a century of written method (the state's first concrete regulations in 1924, AMA since 1950, the Eurocodes since 1975), which is, ironically, exactly what makes its work delegable at all. A method is the part of the job that does not depend on who performs it. The profession wrote its methods down for a hundred years so that any desk would produce the same result. And so, without anyone deciding it, it made its work handable to machines that read.
What the handover leaves behind is the core the methods never covered. An assignment goes out with its purpose, context and constraints; a finished draft comes back; and the engineer's hours concentrate on the three questions that were always the profession's centre of gravity. Does the approach fit this site. Do the assumptions hold. Is a result that looks wrong actually wrong, or the first honest news about the ground. Produced hours stop being the bottleneck. Reviewed judgment becomes the unit firms sell, price, and are scarce in: more alternatives weighed, more decisions per week, the same name under each one.
What it means
For a consultancy, the competition moves from how many hours the firm can produce to how many reviewed, signed decisions it delivers: the firm that reviews and signs three alternatives in the time a competitor produces one owns both the margin and the client conversation.
For a contractor, the same shift lands in the tender room and the site office: the documents that decide bid risk, quantities and change orders can be read completely instead of sampled, and the work preparation that used to be copied from the last project can be built from this activity's documents and this week's conditions.
For both, the procurement test is one question long. Hand any vendor one deliverable your firm actually produces, and judge what comes back: a faster piece of the work, or the finished document, checked and traceable, with the judgment and the signature left where the law puts them.
Software's five public years say the pattern works, what it costs, and where it fails. The built world starts later, with harder documents and stricter rules, and with a century of written method that turns out to have been preparation. What the curve does from here is, for once, this industry's own choice.
Key figures
| Figure | Value | Source |
|---|---|---|
| Global construction productivity growth, two decades | ~1% per year, vs 3.6% in manufacturing | McKinsey Global Institute (2017) |
| Value added per hour since 2000, Sweden | Business sector ~+50%, construction stagnant | SCB via the Productivity Commission (2024) |
| Real repository issues resolved by AI, best published | 1.96% (Oct 2023) → >90% (summer 2026) | SWE-bench / swebench.com |
| Experienced developers with AI assistance, randomised trial | 19% slower, felt ~20% faster | METR (2025) |
| Usable AI task length at 80% vs 50% reliability | About five times shorter | METR (2025) |
| Structural calculation tasks, raw model vs structured system | 0–18% vs 98–100% | Geng et al. (2026) |
| Cost of failed interoperability, US capital facilities | $15.8 billion per year | NIST (2004) |
| The digital model's rank in Swedish standard contracts | "Other document" (AB 04) → own rank proposed (AB 25 draft) | BKK, AB 04 / AB 25 draft |
About this report
Yesper builds AI systems for the category this text describes; that is both why we wrote it and why every claim is sourced to public primary material: launch posts, peer-reviewed studies, state inquiries, standards bodies and regulators, each dated in the text. Calculations of our own are marked as such. Corrections are welcome and will be published in revised editions. First edition, July 2026.
Sources
Yesper is the AI civil engineer for construction and infrastructure. AFRY, COWI, NRC Group and other Nordic firms use it to halve the time on a study, rerun calculations in minutes, and catch errors that would otherwise slip through. Get in touch if you'd like to see what it can do for you.
Book demo