THE FORGE RANKINGS·2026-08-30

Claude Sonnet 4.6

Anthropic
22
RANK of 37 ranked
FERROX INDEX62.6 weighted mean of category positions
GRADEC below the field median
EVIDENCE14 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Not separable from Kimi K2.7 Code at the measured 1.78 point band. The order is printed; the gap is not claimed.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

16of 37 measured
42.9
13.1median 38.963.0

Coding

31of 37 measured
38.3
22.4median 50.070.7

Reasoning

17of 37 measured
65.2
18.1median 62.893.0

Preference

6of 15 measured
73.5
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

ALE-Bench

13/35
1327 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 medium

APEX Agents

18/28
23.7% best 45.0% · Claude Fable 5
APEX CC-BY-4.0 high

LMArena Agent

16/21
0.017 best 0.127 · Claude Opus 5
LMArena CC-BY-4.0

FrontierMath

11/21
32.4% best 47.6% · GPT-5.4 (2026-03-…
Epoch AI CC-BY-4.0 16k

SWE-bench Verified (Epoch's own run)

8/21
75.2% best 83.5% · Claude Opus 4.7
Epoch AI CC-BY-4.0

LMArena Text (style-controlled)

7/20
1472 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

DeepSWE

16/17
29.9% best 72.8% · Claude Opus 5
deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent high

FrontierCode

16/17
24.3% best 53.5% · Claude Fable 5
cognition.com CC-BY-4.0 claude-code max

LMArena WebDev

6/16
1522 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

OSWorld

3/6
72.1% best 75.2% · MiniMax-M3
OSWorld (XLANG Lab) none stated General model 100 steps

Not measured

18/32
  • Humanity's Last Exam
  • METR time horizons
  • Terminal-Bench 4.0
  • Aider Polyglot
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 23.7% 17.6 to 29.8 APEX CC-BY-4.0 / high source
LMArena Agent agents 0.017 18.3 to 26.7 LMArena CC-BY-4.0 source
OSWorld agents 72.1% not published OSWorld (XLANG Lab) none stated General model / 100 steps source
Terminal-Bench 2 agents 53.4% 47.9 to 58.9 tbench.ai CC-BY-4.0 Simplai Agent source
ALE-Bench coding 1327 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 / medium source
DeepSWE coding 29.9% 25.8 to 34.0 deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent / high source
FrontierCode coding 24.3% not published cognition.com CC-BY-4.0 claude-code / max source
SWE-bench Verified (Epoch's own run) coding 75.2% 71.4 to 79.1 Epoch AI CC-BY-4.0 source
LMArena Text (style-controlled) preference 1472 81.2 to 82.3 LMArena CC-BY-4.0 source
LMArena WebDev preference 1522 64.6 to 66.0 LMArena CC-BY-4.0 source
ARC-AGI-2 reasoning 58.3% not published ARC Prize CC-BY-4.0 / max source
FrontierMath reasoning 32.4% 26.9 to 37.9 Epoch AI CC-BY-4.0 / 16k source
GPQA Diamond (Epoch's own run) reasoning 83.3% 78.1 to 88.5 Epoch AI CC-BY-4.0 / high source
OTIS Mock AIME 2024-2025 reasoning 75.6% 62.9 to 88.3 Epoch AI CC-BY-4.0 / high source

Configurations rolled up: 10. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: claude-code, general-model, mini-swe-agent, simplai-agent.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.