THE FORGE RANKINGS·2026-08-30

GPT-5.4 (2026-03-05)

OpenAI
7
RANK of 37 ranked
FERROX INDEX85.5 weighted mean of category positions
GRADEA front rank
EVIDENCE11 / 32 benchmarks measured
QUALITYFULL every category rests on independent benchmarks

Not separable from Claude Opus 4.8 and Claude Opus 4.7 at the measured 1.78 point band. The order is printed; the gap is not claimed.

No measurement in Preference. That weight was redistributed, which is the same as assuming this model would have scored its own average there. It is an assumption, not a measurement.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

1of 37 measured
63.0
13.1median 38.963.0

Coding

11of 37 measured
53.9
22.4median 50.070.7

Reasoning

16of 37 measured
66.0
18.1median 62.893.0

Preference

no qualifying measurement
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

FrontierMath

1/21
47.6% best 47.6% · GPT-5.4 (2026-03-…
Epoch AI CC-BY-4.0 xhigh

SWE-bench Verified (Epoch's own run)

3/21
76.9% best 83.5% · Claude Opus 4.7
Epoch AI CC-BY-4.0 high

DeepSWE

11/17
51.8% best 72.8% · Claude Opus 5
deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent xhigh

Humanity's Last Exam

3/17
36.2% best 46.4% · Gemini 3.1 Pro Pr…
Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0 xhigh

Terminal-Bench 2

2/17
81.8% best 82.2% · GPT-5.5
tbench.ai CC-BY-4.0 ForgeCode

Not measured

21/32
  • LMArena Agent
  • LMArena Text (style-controlled)
  • FrontierCode
  • LMArena WebDev
  • Terminal-Bench 4.0
  • Aider Polyglot
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 34.9% 27.3 to 42.5 APEX CC-BY-4.0 source
METR time horizons agents 5.7h 64.9 to 82.5 metr.org CC-BY-4.0 / xhigh source
Terminal-Bench 2 agents 81.8% 77.9 to 85.7 tbench.ai CC-BY-4.0 ForgeCode source
ALE-Bench coding 1521 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 / medium source
DeepSWE coding 51.8% 50.3 to 53.3 deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent / xhigh source
SWE-bench Verified (Epoch's own run) coding 76.9% 73.1 to 80.6 Epoch AI CC-BY-4.0 / high source
ARC-AGI-2 reasoning 67.5% not published ARC Prize CC-BY-4.0 / high source
FrontierMath reasoning 47.6% 41.9 to 53.3 Epoch AI CC-BY-4.0 / xhigh source
GPQA Diamond (Epoch's own run) reasoning 88.9% 84.5 to 93.3 Epoch AI CC-BY-4.0 / medium source
Humanity's Last Exam reasoning 36.2% 32.6 to 39.9 Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0 / xhigh source
OTIS Mock AIME 2024-2025 reasoning 95.3% 89.0 to 100.0 Epoch AI CC-BY-4.0 / xhigh source

Configurations rolled up: 8. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: forgecode, mini-swe-agent.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.