Claim · agents

Claude Sonnet 4.5on Terminal-Bench 2

40.1 to 42.8 the range across every scoreable measurement we hold
2.7 points between the highest and the lowest
42.6 the number the board uses: the lower median, a real row with its own provenance
5 declared configurations, from 2 runners
0 / 5 rows whose cited page still shows this figure, refetched and checked by us

2 independent runners have measured this, which is rare: on this snapshot only 12 of 321 claim pages can say it.

5 of 5 cited pages no longer show the figure we publish. The page was refetched and split at its own structural boundaries, and no block that names this model carries this number. That does not make the measurement wrong: a leaderboard is a living page and a figure recorded when it was there can be gone by the time you follow the link. It does mean the citation, as printed, no longer supports the number, and this is the only place we know of that says so about its own sources.

What the vendor said, and what someone else measured

Matched on which variant of the benchmark was sat, and on nothing else: the vendor ran it on their harness and the independent party on theirs, and that difference is the comparison rather than an obstacle to it. A vendor number with no independent run of the same variant carries no delta, because a number that cannot be compared honestly must not be compared.

VariantVendor claimsIndependently measuredDifference
as published, no variant stated 46.5 42.6 from 5 runs +3.9

Every measurement

12 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
46.5 / 100
as published: 46.5%
UNATTRIBUTED
shown, never scored
harness CAMEL-AI Terminal-Bench
via Epoch AI
2025-12-24
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
43.1 / 100
as published: 43.1%
UNATTRIBUTED
shown, never scored
harness Goose Terminal-Bench
via Epoch AI
2025-12-11
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
42.8 / 100
as published: 42.8%
UNATTRIBUTED
shown, never scored
harness Terminus 2 Terminal-Bench
via Epoch AI
2025-10-31
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
42.8 / 100
as published: 42.8%
INDEPENDENT
counts toward the ranking
harness Terminus 2 Terminal-Bench v2 Leaderboard
via Epoch AI
2025-10-31
CC-BY-4.0
checked: this figure is not on the page today2026-08-30
Terminal-Bench v2 Leaderboard
42.7 / 100
as published: 42.7%
INDEPENDENT
counts toward the ranking
harness MAYA-V2 tbench.ai
via Epoch AI
2026-01-04
CC-BY-4.0
checked: this figure is not on the page today2026-08-30
leaderboard row
42.7 / 100
as published: 42.7%
UNATTRIBUTED
shown, never scored
harness MAYA Terminal-Bench
via Epoch AI
2026-01-04
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
42.6 / 100
as published: 42.6%
UNATTRIBUTED
shown, never scored
harness OpenHands Terminal-Bench
via Epoch AI
2025-11-02
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
42.6 / 100
as published: 42.6%
INDEPENDENT
counts toward the ranking
harness OpenHands Terminal-Bench v2 Leaderboard
via Epoch AI
2025-11-02
CC-BY-4.0
checked: this figure is not on the page today2026-08-30
Terminal-Bench v2 Leaderboard
42.5 / 100
as published: 42.5%
UNATTRIBUTED
shown, never scored
harness Mini-SWE-Agent Terminal-Bench
via Epoch AI
2025-11-03
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
42.5 / 100
as published: 42.5%
INDEPENDENT
counts toward the ranking
harness Mini-SWE-Agent Terminal-Bench v2 Leaderboard
via Epoch AI
2025-11-03
CC-BY-4.0
checked: this figure is not on the page today2026-08-30
Terminal-Bench v2 Leaderboard
40.1 / 100
as published: 40.1%
UNATTRIBUTED
shown, never scored
harness Claude Code Terminal-Bench
via Epoch AI
2025-11-04
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
40.1 / 100
as published: 40.1%
INDEPENDENT
counts toward the ranking
harness Claude Code Terminal-Bench v2 Leaderboard
via Epoch AI
2025-11-04
CC-BY-4.0
checked: this figure is not on the page today2026-08-30
Terminal-Bench v2 Leaderboard
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.