AI made developers 19% slower. What that teaches engineering firms.

In July 2025 the research group METR published a randomized trial: experienced developers using AI assistants took 19 percent longer to finish their tasks, while estimating they had been 24 percent faster. The lesson is not that AI fails. It is that the shape of the task decides whether AI helps, and what is worth paying for.

Freshly cast concrete surface with forested hills in the background

METR measured 246 real tasks, not toy problems

METR, a non-profit AI research group, recruited 16 experienced open-source developers and randomized 246 of their real tasks: issues, bugs and features in mature repositories the developers had worked in for an average of five years. On some tasks they could use AI tools, mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet. On the others they could not. No toy problems, no hired novices: their own code, their own backlog.

Before starting, the developers forecast that AI would cut their completion time by 24 percent. The measurement went the other way: with AI allowed, tasks took 19 percent longer. The forecast missed by 43 percentage points, and that gap may be the study's most useful finding. The people closest to the work misjudged not just the size of the effect but its direction.

An earlier trial found the opposite: 55.8 percent faster

Two years earlier, researchers at Microsoft Research and MIT had run a randomized trial that found the opposite. Peng and colleagues gave 95 developers a defined task: implement an HTTP server in JavaScript, with automated tests as the finish line. The group with access to GitHub Copilot completed the task 55.8 percent faster than the control group (p = 0.0017), and succeeded slightly more often: 78 percent against 70. That figure travelled far and became the default expectation for what AI does to knowledge work.

Both trials are careful. Neither is wrong. What differs is the shape of the work.

Study Result Task shape
Peng et al., Microsoft Research and MIT (2023) 55.8% faster Bounded, specified in advance, verified by automated tests
METR (2025), measured 19% slower Open-ended tasks in the developers' own mature codebases
METR (2025), developers' own forecast 24% faster The same tasks, self-assessed before starting

The shape of the task decides the result

In the Microsoft and MIT trial the unit of work was whole and checkable: a defined deliverable with an objective test at the end. In the METR trial AI was an assistant inside open-ended expert work. Developers prompted, waited, read plausible suggestions and corrected them, in codebases they knew better than the model did. Every step left the burden of verification on the human.

That is the pattern, and it has little to do with model quality. When AI delivers a whole piece of work with a cheap, objective check at the end, machine time replaces human time. When AI delivers a stream of suggestions inside an expert's judgment, the expert becomes a reviewer of drafts, and reviewing plausible drafts can cost more than writing from scratch.

The same boundary is visible in engineering documents. In March 2026 the AI company Nomic published AEC-Bench, a benchmark of 196 real tasks from architecture, engineering and construction. On submittal review, checking a contractor's technical submittals against specifications and drawings, the paper reports the best frontier-model score at 23.1 percent. It is one vendor's benchmark, so read it as directional. But the direction is familiar: generic model, open-ended domain task, weak result.

Buy completed deliverables, not assistance

Completed deliverables with verification, not assistance. A chat assistant rolled out to experienced engineers on their own projects reproduces the METR condition: open-ended help inside deep expertise, verification left to the human at every step. The Peng condition looks different in an engineering firm: a bounded deliverable such as a traffic study or a compliance check against a specification, produced whole, checked systematically, then reviewed and signed by the engineer.

The second lesson: measure, do not survey. A 43-point gap between predicted and actual speed means pilot questionnaires mostly measure enthusiasm. Count hours on comparable deliverables, before and after, on real projects. If a vendor's evidence is a satisfaction score, ask for timesheets.

This is also the yardstick we accept for ourselves. Our defensible claims are deliberately narrow, and measured the way this section asks: a whole deliverable that used to take the better part of a fortnight, now produced and verified in a day on a real project, not averaged into a slogan; and one extra design iteration per project at constant fee. The 19 percent study is not an argument against AI in engineering. It is the buying guide.

  1. Joel Becker, Nate Rush, Beth Barnes & David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", METR, arXiv:2507.09089, July 2025.
  2. Sida Peng, Eirini Kalliamvakou, Peter Cihon & Mert Demirer, "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot", Microsoft Research and MIT, arXiv:2302.06590, 2023.
  3. Harsh Mankodiya, Chase Gallik, Theodoros Galanos & Andriy Mulyar, "AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction", arXiv:2603.29199, March 2026.
Benjamin Glaser Co-founder at Yesper. Writes about AI and the industry that builds the world. benjamin@yesper.ai

Yesper is the AI civil engineer for construction and infrastructure. AFRY, COWI, NRC Group and other Nordic firms use it to halve the time on a study, rerun calculations in minutes, and catch errors that would otherwise slip through. Get in touch if you'd like to see what it can do for you.

Book demo