Skip to content
Nornic
The Claim Register

Nornic Register · Living document

AI agents can now work autonomously for days.

Cannot be checked

The most-cited public instrument for agent task length says it stops being reliable at sixteen hours. Past that, there is nothing to check the claim against.

The source

Study
Task-Completion Time Horizons of Frontier AI Models
Publisher
METR
Publication type
Ongoing public measurement, interactive
Sample
Frontier models against METR’s task suite; human baselines from contracted professionals
Fieldwork
Updated 8 May 2026
Funding & conflicts
METR — a non-profit AI evaluation organisation
Verified on
2026-08-09

The arithmetic

METR measures capability as a 50%-time horizon: in their words, "the length of task in our suite (measured by how long it takes a human expert) such that we'd predict with 50% confidence that the model could complete the task."

The ceiling is stated on the page, unprompted: "Measurements above 16 hrs are unreliable with our current task suite."

Sixteen hours is two working days for one person. So the instrument covers roughly the range a vendor demo occupies — and stops right where the more ambitious marketing starts.

This is not a claim that long-running agents fail. It is narrower and more awkward: above about a working day, there is no accepted public measurement that would tell you either way. Absence of an instrument is not evidence of absence.

Claims in the single-digit hours are inside the measurable range and are not disputed here. The verdict applies to "days", not to "hours".

The baselines are worth knowing when reading any horizon number: humans contracted by METR are professionals in software engineering, ML or security, averaging about five years of relevant experience, and task durations come from the geometric mean of their successful completion times. The comparison is against competent practitioners, not novices.

Nornic makes no claim about how long an agent can run unattended, because there is nothing to base one on. That is the entry.

What would change this verdict

A published task suite with reliable measurements past sixteen hours, and a horizon figure measured on it. METR extending its own suite would do it; so would an independent instrument with an inspectable method.

Sources