Claim · agents

Claude 3.7 Sonneton METR time horizons

50.0 to 50.0 the range across every scoreable measurement we hold
0.0 points between the highest and the lowest
50.0 the number the board uses: the lower median, a real row with its own provenance
1 declared configuration, from 1 runner
1 / 1 rows whose cited page still shows this figure, refetched and checked by us

All 1 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.

What the vendor said, and what someone else measured

Matched on which variant of the benchmark was sat, and on nothing else: the vendor ran it on their harness and the independent party on theirs, and that difference is the comparison rather than an obstacle to it. A vendor number with no independent run of the same variant carries no delta, because a number that cannot be compared honestly must not be compared.

VariantVendor claimsIndependently measuredDifference
as published, no variant stated 50.9 50.0 from 1 run +0.9

Every measurement

2 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
50.9 / 100
as published: 1.0h
UNATTRIBUTED
shown, never scored
not stated by the source METR
via Epoch AI
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
50.0 / 100
as published: 56m
INDEPENDENT
counts toward the ranking
effort 16k METR - Measuring AI Ability to Complete Long Tasks
via Epoch AI
CC-BY-4.0
checked: this figure appears in the page’s row for this model2026-08-30Depiction of the process of computing the time horizon. For example, Claude 3.7 Sonnet (the right-most model, represented in the darkest green) has a time horiz
METR - Measuring AI Ability to Complete Long Tasks
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.