All 1 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.
Matched on which variant of the benchmark was sat, and on nothing else: the vendor ran it on their harness and the independent party on theirs, and that difference is the comparison rather than an obstacle to it. A vendor number with no independent run of the same variant carries no delta, because a number that cannot be compared honestly must not be compared.
| Variant | Vendor claims | Independently measured | Difference |
|---|---|---|---|
| as published, no variant stated | 60.0 | 55.0 from 1 run | +5.0 |
2 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.
| Score | Provenance | Configuration | Who ran it | Where to check it |
|---|---|---|---|---|
| 60.0 / 100
as published: 60.0% |
CLAIMED
shown, never scored |
not stated by the source | Opus 4.5 System Card via Epoch AI CC-BY-4.0 |
not checked: the page is a format we cannot read2026-08-30 Opus 4.5 System Card |
| 55.0 / 100
as published: 55.0% |
INDEPENDENT
counts toward the ranking |
not stated by the source | Cybench leaderboard via Epoch AI CC-BY-4.0 |
checked: this figure is not on the page today2026-08-30 Cybench leaderboard² Results from the Claude Sonnet 4.5 System Card on a subset of 37 problems. Percentages are estimates based on probability of success on 1 trial. |
Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.