Skip to content
Nornic
The Claim Register

Nornic Register · Living document

AI makes developers 19% slower.

Over-extended

A real result from sixteen experienced maintainers working on code they already knew, quoted as a fact about the profession.

The source

Study
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
Publisher
METR
Publication type
Randomised controlled trial, published research
Sample
16 developers, 246 tasks
Fieldwork
February – June 2025
Funding & conflicts
METR — a non-profit AI evaluation organisation
Verified on
2026-08-09

The arithmetic

The study is careful and its authors are unusually candid about its limits. The problem is not the study; it is the citation.

The measured result stands: allowed AI tools, the developers took 19% longer, with a confidence interval of +2% to +39%. Nothing below disputes it.

n = 16, on 246 issues in repositories the participants already maintained — averaging over 22k stars and a million lines. That supports a claim about experienced maintainers on familiar code. Deep familiarity with a codebase is the condition under which AI assistance has least to add, so it is close to the least favourable case, not a representative one.

The authors say so themselves: "We do not claim that our developers or repositories represent a majority or plurality of software development work." The headline drops the condition; the paper never did.

The durable finding is not the 19%. Beforehand the developers expected AI to speed them up by 24%. Afterwards, having been measurably slower, they still believed they had been 20% faster. Experienced people could not tell which direction they had moved in — and that survives every caveat here.

The update almost nobody cites: in February 2026 METR reported that its follow-up — 10 developers from the original plus 47 new ones, running from August 2025 — had a selection problem it named itself, "systematically missing developers who have the most optimistic expectations about AI's value" and "systematically missing tasks which have high expected uplift from AI".

Stated precisely, because precision is the whole point: the follow-up was redesigned, not retracted. METR's post is titled "We are Changing our Developer Productivity Experiment Design", and the later data ran the other way — an estimated speedup of 18%, confidence interval −38% to +9%, which includes zero and so settles nothing on its own.

A research team publishing the reason to doubt its own most-quoted number is the behaviour worth copying here. The number is not.

What would change this verdict

A replication at meaningful scale, on a population that is not exclusively expert maintainers working in code they wrote, with the selection problem METR identified addressed. METR is currently building it.

Sources