THE FORGE RANKINGS·2026-08-30

Claude Opus 5

Anthropic
1
RANK of 37 ranked
FERROX INDEX99.5 weighted mean of category positions
GRADEA+ top of the measured field
EVIDENCE9 / 32 benchmarks measured
QUALITYFULL every category rests on independent benchmarks

No measurement in Preference. That weight was redistributed, which is the same as assuming this model would have scored its own average there. It is an assumption, not a measurement.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

4of 37 measured
51.5
13.1median 38.963.0

Coding

2of 37 measured
62.7
22.4median 50.070.7

Reasoning

1of 37 measured
93.0
18.1median 62.893.0

Preference

no qualifying measurement
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

ALE-Bench

2/35
2165 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 high

APEX Agents

2/28
43.5% best 45.0% · Claude Fable 5
APEX CC-BY-4.0 max

FrontierCode

2/17
53.4% best 53.5% · Claude Fable 5
cognition.com CC-BY-4.0 claude-code max

Terminal-Bench 4.0

1/10
51.8% best 51.8% · Claude Opus 5
Terminal-Bench Apache-2.0 claude-code max

Not measured

23/32
  • FrontierMath
  • SWE-bench Verified (Epoch's own run)
  • LMArena Text (style-controlled)
  • Humanity's Last Exam
  • Terminal-Bench 2
  • LMArena WebDev
  • METR time horizons
  • Aider Polyglot
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 43.5% 35.3 to 51.7 APEX CC-BY-4.0 / max source
LMArena Agent agents 0.127 52.5 to 65.6 LMArena CC-BY-4.0 source
Terminal-Bench 4.0 agents 51.8% 48.4 to 55.2 Terminal-Bench Apache-2.0 claude-code / max source
ALE-Bench coding 2165 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 / high source
DeepSWE coding 72.8% 70.9 to 74.8 deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent / high source
FrontierCode coding 53.4% not published cognition.com CC-BY-4.0 claude-code / max source
ARC-AGI-2 reasoning 88.3% not published ARC Prize CC-BY-4.0 / max source
GPQA Diamond (Epoch's own run) reasoning 92.9% 89.3 to 96.5 Epoch AI CC-BY-4.0 source
OTIS Mock AIME 2024-2025 reasoning 97.8% 93.4 to 100.0 Epoch AI CC-BY-4.0 source

Configurations rolled up: 10. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: claude-code, mini-swe-agent.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.