THE FORGE RANKINGS·2026-08-30

GPT-OSS 120B

OpenAI
35
RANK of 37 ranked
FERROX INDEX24.1 weighted mean of category positions
GRADEE bottom of the measured field
EVIDENCE8 / 32 benchmarks measured
QUALITYFULL every category rests on independent benchmarks

No measurement in Preference. That weight was redistributed, which is the same as assuming this model would have scored its own average there. It is an assumption, not a measurement.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

36of 37 measured
21.8
13.1median 38.963.0

Coding

35of 37 measured
29.1
22.4median 50.070.7

Reasoning

7of 37 measured
82.3
18.1median 62.893.0

Preference

no qualifying measurement
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

GPQA Diamond (Epoch's own run)

35/36
75.8% best 94.1% · Gemini 3.1 Pro Pr…
Epoch AI CC-BY-4.0 high

OTIS Mock AIME 2024-2025

19/36
88.9% best 99.7% · Claude Fable 5
Epoch AI CC-BY-4.0 high

ALE-Bench

35/35
576 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0

APEX Agents

28/28
4.7% best 45.0% · Claude Fable 5
APEX CC-BY-4.0

LMArena Text (style-controlled)

20/20
1352 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

METR time horizons

11/11
42m best 6.4h · Gemini 3.1 Pro Pr…
METR - Measuring AI Ability to Complete Long Tasks CC-BY-4.0

Not measured

24/32
  • ARC-AGI-2
  • LMArena Agent
  • FrontierMath
  • SWE-bench Verified (Epoch's own run)
  • DeepSWE
  • FrontierCode
  • Humanity's Last Exam
  • LMArena WebDev
  • Terminal-Bench 4.0
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 4.7% 2.0 to 7.4 APEX CC-BY-4.0 source
METR time horizons agents 42m 36.4 to 54.6 METR - Measuring AI Ability to Complete Long Tasks CC-BY-4.0 source
Terminal-Bench 2 agents 14.2% 9.7 to 18.7 Terminal-Bench v2 Leaderboard CC-BY-4.0 Mini-SWE-Agent source
Aider Polyglot coding 41.8% not published Aider LLM Leaderboards CC-BY-4.0 diff / high source
ALE-Bench coding 576 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 source
LMArena Text (style-controlled) preference 1352 64.0 to 65.3 LMArena CC-BY-4.0 source
GPQA Diamond (Epoch's own run) reasoning 75.8% 70.4 to 81.1 Epoch AI CC-BY-4.0 / high source
OTIS Mock AIME 2024-2025 reasoning 88.9% 80.1 to 97.6 Epoch AI CC-BY-4.0 / high source

Configurations rolled up: 5. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: diff, mini-swe-agent, terminus-2.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.