THE FORGE RANKINGS·2026-08-30

Claude 3.7 Sonnet

Anthropic
36
RANK of 37 ranked
FERROX INDEX23.3 weighted mean of category positions
GRADEE bottom of the measured field
EVIDENCE11 / 32 benchmarks measured
QUALITYFULL every category rests on independent benchmarks

No measurement in Preference. That weight was redistributed, which is the same as assuming this model would have scored its own average there. It is an assumption, not a measurement.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

26of 37 measured
35.2
13.1median 38.963.0

Coding

3of 37 measured
60.7
22.4median 50.070.7

Reasoning

35of 37 measured
27.5
18.1median 62.893.0

Preference

no qualifying measurement
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

SWE-bench Verified (Epoch's own run)

20/21
61.0% best 83.5% · Claude Opus 4.7
Epoch AI CC-BY-4.0

LMArena Text (style-controlled)

19/20
1372 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

Humanity's Last Exam

12/17
8.0% best 46.4% · Gemini 3.1 Pro Pr…
Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0

Cybench

5/5
20.0% best 55.0% · Claude Sonnet 4.5
Cybench leaderboard CC-BY-4.0

Not measured

21/32
  • ALE-Bench
  • APEX Agents
  • LMArena Agent
  • DeepSWE
  • FrontierCode
  • Terminal-Bench 2
  • LMArena WebDev
  • Terminal-Bench 4.0
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
Cybench agents 20.0% not published Cybench leaderboard CC-BY-4.0 source
METR time horizons agents 56m 41.9 to 56.3 METR - Measuring AI Ability to Complete Long Tasks CC-BY-4.0 / 16k source
OSWorld agents 35.6% not published OSWorld (XLANG Lab) none stated General model / 100 steps source
Aider Polyglot coding 60.4% not published aider.chat CC-BY-4.0 diff source
SWE-bench Verified (Epoch's own run) coding 61.0% 56.6 to 65.3 Epoch AI CC-BY-4.0 source
LMArena Text (style-controlled) preference 1372 66.8 to 68.0 LMArena CC-BY-4.0 source
ARC-AGI-2 reasoning 0.4% not published ARC Prize CC-BY-4.0 / 1k source
FrontierMath reasoning 3.1% 1.1 to 5.1 Epoch AI CC-BY-4.0 / 64k source
GPQA Diamond (Epoch's own run) reasoning 76.8% 70.9 to 82.7 Epoch AI CC-BY-4.0 / 16k source
Humanity's Last Exam reasoning 8.0% 5.9 to 10.1 Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0 source
OTIS Mock AIME 2024-2025 reasoning 46.7% 31.9 to 61.4 Epoch AI CC-BY-4.0 / 16k source

Configurations rolled up: 11. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: diff, general-model.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.