Claim · reasoning

Claude Sonnet 4.5on Humanity's Last Exam

7.5 to 13.7 the range across every scoreable measurement we hold
6.2 points between the highest and the lowest
7.5 the number the board uses: the lower median, a real row with its own provenance
1 declared configuration, from 1 runner

All 2 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.

Two measurements, one declared configuration, +6.2 points apart. 2 rows here agree on everything the record states, the system identity included, and disagree on the number: 13.7, 7.5. Something moved the score and no source we hold says what. The board resolves it by taking the lower median, which is a rule and not an explanation, so it is named here rather than quietly settled. This happens in 14 of 321 claim pages.

Every measurement

2 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
13.7 / 100
as published: 13.7%
INDEPENDENT
counts toward the ranking
not stated by the source Humanity’s Last Exam (CAIS / Scale AI)
via Epoch AI
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
7.5 / 100
as published: 7.5%
INDEPENDENT
counts toward the ranking
not stated by the source Humanity’s Last Exam (CAIS / Scale AI)
via Epoch AI
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.