Claim · reasoning

Claude 3.7 Sonneton GPQA Diamond (Epoch's own run)

66.0 to 78.5 the range across every scoreable measurement we hold
12.5 points between the highest and the lowest
76.8 the number the board uses: the lower median, a real row with its own provenance
4 declared configurations, from 1 runner

All 4 scoreable measurements here come from ONE runner, so this is not a reconciliation and is not presented as one. What it shows is how far a single runner's number moves when the model is run differently.

Every measurement

4 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
78.5 / 100
as published: 78.5%
INDEPENDENT
counts toward the ranking
effort 64k Epoch AI
2025-05-26T22:53:45.927Z
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
76.8 / 100
as published: 76.8%
INDEPENDENT
counts toward the ranking
effort 32k Epoch AI
2025-03-10T22:05:35.890Z
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
76.8 / 100
as published: 76.8%
INDEPENDENT
counts toward the ranking
effort 16k Epoch AI
2025-02-26T16:56:43.203Z
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
66.0 / 100
as published: 66.0%
INDEPENDENT
counts toward the ranking
not stated by the source Epoch AI
2025-02-24T19:14:07.847Z
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.