THE FORGE RANKINGS·2026-08-30

Grok 4.5

xAI
11
RANK of 37 ranked
FERROX INDEX81.0 weighted mean of category positions
GRADEA front rank
EVIDENCE11 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Not separable from GPT-5.6 Terra at the measured 1.78 point band. The order is printed; the gap is not claimed.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

32of 37 measured
28.3
13.1median 38.963.0

Coding

26of 37 measured
44.5
22.4median 50.070.7

Reasoning

8of 37 measured
81.3
18.1median 62.893.0

Preference

4of 15 measured
75.5
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

GPQA Diamond (Epoch's own run)

2/36
93.4% best 94.1% · Gemini 3.1 Pro Pr…
Epoch AI CC-BY-4.0 high

OTIS Mock AIME 2024-2025

4/36
97.8% best 99.7% · Claude Fable 5
Epoch AI CC-BY-4.0 high

ALE-Bench

16/35
1309 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 high

APEX Agents

10/28
34.2% best 45.0% · Claude Fable 5
APEX CC-BY-4.0

LMArena Agent

8/21
0.065 best 0.127 · Claude Opus 5
LMArena CC-BY-4.0

LMArena Text (style-controlled)

8/20
1470 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

DeepSWE

9/17
53.8% best 72.8% · Claude Opus 5
deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent high

FrontierCode

9/17
42.4% best 53.5% · Claude Fable 5
cognition.com CC-BY-4.0 grok-build high

LMArena WebDev

3/16
1556 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

Terminal-Bench 4.0

10/10
12.4% best 51.8% · Claude Opus 5
Terminal-Bench Apache-2.0 grok-build none

Not measured

21/32
  • FrontierMath
  • SWE-bench Verified (Epoch's own run)
  • Humanity's Last Exam
  • Terminal-Bench 2
  • METR time horizons
  • Aider Polyglot
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 34.2% 26.8 to 41.6 APEX CC-BY-4.0 source
LMArena Agent agents 0.065 34.3 to 42.2 LMArena CC-BY-4.0 source
Terminal-Bench 4.0 agents 12.4% 9.8 to 15.0 Terminal-Bench Apache-2.0 grok-build / none source
ALE-Bench coding 1309 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 / high source
DeepSWE coding 53.8% 51.5 to 56.0 deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent / high source
FrontierCode coding 42.4% not published cognition.com CC-BY-4.0 grok-build / high source
LMArena Text (style-controlled) preference 1470 80.6 to 82.1 LMArena CC-BY-4.0 source
LMArena WebDev preference 1556 68.5 to 70.6 LMArena CC-BY-4.0 source
ARC-AGI-2 reasoning 52.6% not published ARC Prize CC-BY-4.0 / high source
GPQA Diamond (Epoch's own run) reasoning 93.4% 90.6 to 96.3 Epoch AI CC-BY-4.0 / high source
OTIS Mock AIME 2024-2025 reasoning 97.8% 95.3 to 100.0 Epoch AI CC-BY-4.0 / high source

Configurations rolled up: 7. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: grok-build, mini-swe-agent.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.