In July 2025 the research group METR published a randomized trial: experienced developers using AI assistants took 19 percent longer to finish their tasks, while estimating they had been 24 percent faster. The lesson is not that AI fails. It is that the shape of the task decides whether AI helps, and what is worth paying for.
The study
METR, a non-profit AI research group, recruited 16 experienced open-source developers and randomized 246 of their real tasks: issues, bugs and features in mature repositories the developers had worked in for an average of five years. On some tasks they could use AI tools, mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet. On the others they could not. No toy problems, no hired novices: their own code, their own backlog.
Before starting, the developers forecast that AI would cut their completion time by 24 percent. The measurement went the other way: with AI allowed, tasks took 19 percent longer. The forecast missed by 43 percentage points, and that gap may be the study's most useful finding. The people closest to the work misjudged not just the size of the effect but its direction.
The other study
Two years earlier, researchers at Microsoft Research and MIT had run a randomized trial that found the opposite. Peng and colleagues gave 95 developers a defined task: implement an HTTP server in JavaScript, with automated tests as the finish line. The group with access to GitHub Copilot completed the task 55.8 percent faster than the control group (p = 0.0017), and succeeded slightly more often: 78 percent against 70. That figure travelled far and became the default expectation for what AI does to knowledge work.
Both trials are careful. Neither is wrong. What differs is the shape of the work.
| Study | Result | Task shape |
|---|---|---|
| Peng et al., Microsoft Research and MIT (2023) | 55.8% faster | Bounded, specified in advance, verified by automated tests |
| METR (2025), measured | 19% slower | Open-ended tasks in the developers' own mature codebases |
| METR (2025), developers' own forecast | 24% faster | The same tasks, self-assessed before starting |
The difference
In the Microsoft and MIT trial the unit of work was whole and checkable: a defined deliverable with an objective test at the end. In the METR trial AI was an assistant inside open-ended expert work. Developers prompted, waited, read plausible suggestions and corrected them, in codebases they knew better than the model did. Every step left the burden of verification on the human.
That is the pattern, and it has little to do with model quality. When AI delivers a whole piece of work with a cheap, objective check at the end, machine time replaces human time. When AI delivers a stream of suggestions inside an expert's judgment, the expert becomes a reviewer of drafts, and reviewing plausible drafts can cost more than writing from scratch.
The same boundary is visible in engineering documents. In March 2026 the AI company Nomic published AEC-Bench, a benchmark of 196 real tasks from architecture, engineering and construction. On submittal review, checking a contractor's technical submittals against specifications and drawings, the paper reports the best frontier-model score at 23.1 percent. It is one vendor's benchmark, so read it as directional. But the direction is familiar: generic model, open-ended domain task, weak result.
The lesson
Completed deliverables with verification, not assistance. A chat assistant rolled out to experienced engineers on their own projects reproduces the METR condition: open-ended help inside deep expertise, verification left to the human at every step. The Peng condition looks different in an engineering firm: a bounded deliverable such as a traffic study or a compliance check against a specification, produced whole, checked systematically, then reviewed and signed by the engineer.
The second lesson: measure, do not survey. A 43-point gap between predicted and actual speed means pilot questionnaires mostly measure enthusiasm. Count hours on comparable deliverables, before and after, on real projects. If a vendor's evidence is a satisfaction score, ask for timesheets.
This is also the yardstick we accept for ourselves. Our defensible claims are deliberately narrow, and measured the way this section asks: a whole deliverable that used to take the better part of a fortnight, now produced and verified in a day on a real project, not averaged into a slogan; and one extra design iteration per project at constant fee. The 19 percent study is not an argument against AI in engineering. It is the buying guide.
Sources
Yesper is the AI civil engineer for construction and infrastructure. AFRY, COWI, NRC Group and other Nordic firms use it to halve the time on a study, rerun calculations in minutes, and catch errors that would otherwise slip through. Get in touch if you'd like to see what it can do for you.
Book demoNot ready for a demo? Get the next piece in your inbox.