THE FORGE / RANKINGS

Benchmark report

This board ranks public measurements, not models. Every number names the benchmark that produced it, the evaluator who ran that benchmark, the source it came through and the licence it was published under. Vendor self-reports are recorded and never scored. Ferrox has run none of these evaluations itself; where a measurement cannot be attributed to a named evaluator, the row says so and does not enter a score.

Snapshot
2026-08-30-62e4e85f043b
Source manifest
dc46131315046f1e
As of
2026-08-30
Rubric
v1.1.0
Stat contract
v1.0.0
UNCALIBRATED

Anchors and thresholds are placeholders until the 90-day backfill study runs. Every published number must carry this label.

Read the ordering, not the absolute value. The rubric is uncalibrated, so the gap between two rows is far more trustworthy than the number printed on either of them.

01 At a glance

2343

Observations ingested

every row read from every source

2125

Entered a score

90.7% of what was read

1147

Systems identified

model, harness and effort tuples

37

Models ranked

of 610 models seen

2

Systems ranked

under the strict tuple entity

32

Benchmarks in the score

each one declared in the stat contract

02 How to read this report

Two words this report cannot do without

Harness
The scaffolding a model is run inside on an agent benchmark: the loop that gives it tools, feeds back results and decides when to stop. The same model scores very differently under different harnesses, so a measurement is only meaningful if the harness is named. Where a publisher does not name one, this report says "as published" rather than inventing one.
Effort
The reasoning setting a model was run at, where the publisher exposes one, for example low, high or max. It changes both cost and score, so it is part of what was measured.

Two entities, and they are not the same claim

A system is one exact combination of model, harness and effort. It is the rigorous unit: two measurements are joined only when all three match, so a run whose harness was never published is never merged with one whose harness was.

A model is that model across every configuration measured. Rolling up to it is a declared operation with a stated rule, not a guess, and every model row names how many configurations sit behind it and which harnesses they used.

The rollup rule is stated, not implied: median observed configuration per benchmark (lower median, always a real measurement). Taking the best configuration instead would reward a model for having been measured more often, because the maximum of more draws is higher whether or not the model is better.

What has to be true before a score is published

A Ferrox Index score requires at least 2 qualifying benchmarks in each of agents, coding, reasoning. Systems below that bar appear on the category boards with their evidence count shown, and have no overall score.

The evidence flag on every row

What the interval does and does not say

The interval printed beside an index is built only from the benchmarks that published one. Five of the benchmarks scored here publish no uncertainty at all, so for most rows the interval is a lower bound: the true uncertainty is at least this wide and can only be wider. Rows in that state are marked "at least". Publishing no interval at all would have thrown away the real information the other sources did provide.

How rows are ordered

Rows are ordered by the point estimate. The interval is printed beside it and is never used as the sort key. Sorting by the top of the interval instead would make a wider interval a rank bonus, and would reward a source for publishing a vaguer number.

03 The board

Models, rolled up across every configuration measured. 37 of 610 models clear the entry requirement.

#ModelOrganisationWeights Ferrox IndexOld composite BenchesConfigsRule costEvidence
1 Claude Opus 5 imputed: preference Anthropic closed 99.5 66.7 9
of 32
10
claude-code, mini-swe-agent
1.17 FULL
2 GPT-5.6 Sol imputed: preference not separable from 3 OpenAI closed 95.5 61.3 9
of 32
13
codex, mini-swe-agent
4.12 FULL
3 Claude Fable 5 not separable from 2, 4 Anthropic closed 94.8 65.8 11
of 32
13
claude-code, mini-swe-agent
1.45 THIN
4 Kimi K3 imputed: preference not separable from 3 Moonshot open 93.7 56.5 8
of 32
6
mini-swe-agent
3.50 FULL
5 Grok 4.6 imputed: preference xAI closed 92.8 53.5 9
of 32
11
grok-build, mini-swe-agent
2.58 FULL
6 Claude Opus 4.8 not separable from 7 Anthropic closed 87.1 51.3 12
of 32
12
claude-code, mini-swe-agent
6.63 THIN
7 GPT-5.4 (2026-03-05) imputed: preference not separable from 6, 9 OpenAI unknown 85.5 60.8 11
of 32
8
forgecode, mini-swe-agent
3.38 FULL
8 GLM-5.3 imputed: preference not separable from 9 Z.ai (Zhipu AI) unknown 84.9 57.0 6
of 32
4
claude-code, mini-swe-agent
0.00 FULL
9 Claude Opus 4.7 not separable from 7, 8 Anthropic closed 84.0 57.6 13
of 32
8
claude-code, wozcode
5.75 THIN
10 GPT-5.5 not separable from 12 OpenAI closed 81.6 58.0 11
of 32
17
capy, clnkr, codex +3
8.98 THIN
11 Grok 4.5 not separable from 12 xAI closed 81.0 51.1 11
of 32
7
grok-build, mini-swe-agent
0.00 THIN
12 GPT-5.6 Terra imputed: preference not separable from 10, 11 OpenAI closed 80.1 48.7 8
of 32
12
codex, mini-swe-agent
12.36 FULL
13 Claude Sonnet 5 imputed: preference Anthropic closed 74.2 48.3 8
of 32
11
claude-code, mini-swe-agent
10.19 FULL
14 GLM-5.2 imputed: preference not separable from 15, 17 Z.ai (Zhipu AI) unknown 71.1 43.6 9
of 32
8
mini-swe-agent
5.84 FULL
15 Claude Opus 4.6 not separable from 14, 16 Anthropic closed 70.8 52.3 13
of 32
11
capy, claude-code, droid +3
3.81 THIN
16 GPT-5.6 Luna imputed: preference not separable from 15, 17 OpenAI closed 70.8 40.0 8
of 32
12
codex, mini-swe-agent
23.53 FULL
17 Gemini 3.1 Pro Preview not separable from 14, 16 Google DeepMind closed 69.9 48.4 13
of 32
7
forge-code, forgecode, gemini-cli +2
0.55 THIN
18 Kimi K2.6 not separable from 19, 20, 21 Moonshot unknown 65.9 56.1 10
of 32
2
general-model
0.00 THIN
19 GPT-5.2 (2025-12-11) imputed: preference not separable from 18 OpenAI unknown 65.5 50.5 9
of 32
8
codex-cli, droid
4.82 FULL
20 Gemini 3 Pro Preview imputed: preference not separable from 18 Google DeepMind closed 65.3 51.4 9
of 32
5
codebrain-1-5, droid, ii-agent +1
1.10 FULL
21 Gemini 3 Flash Preview imputed: preference not separable from 18 Google DeepMind closed 65.0 47.6 8
of 32
4
gemini-cli, junie-cli
13.14 FULL
22 Claude Sonnet 4.6 not separable from 23 Anthropic closed 62.6 50.2 14
of 32
10
claude-code, general-model, mini-swe-agent +1
5.46 THIN
23 Kimi K2.7 Code imputed: preference not separable from 22, 24 Moonshot open 61.5 44.7 8
of 32
3
mini-swe-agent
0.00 FULL
24 Claude 4.5 Opus not separable from 23 Anthropic closed 61.4 50.7 12
of 32
7
droid, letta-code
7.30 THIN
25 GLM 5.1 Z.ai (Zhipu AI) unknown 58.7 53.3 13
of 32
7
ascend-a3-dual-node, ascend-a5-w4a4c8, gpu-baseline;-10-groups-over-3-rounds +3
10.17 THIN
26 GPT-5.1 (2025-11-13) imputed: preference OpenAI closed 47.4 38.7 9
of 32
6
terminus-2
9.46 FULL
27 GPT-5 (2025-08-07) imputed: preference OpenAI unknown 45.2 46.8 11
of 32
12
codex-cli, diff, mini-swe-agent +2
4.60 FULL
28 Kimi K2.5 imputed: preference Moonshot unknown 43.4 36.3 6
of 32
2
general-model
0.00 FULL
29 o3 (2025-04-16) imputed: preference not separable from 30 OpenAI unknown 42.2 43.3 11
of 32
7
diff
2.07 FULL
30 MiniMax-M3 not separable from 29 MiniMax open 40.6 45.6 9
of 32
5
general-model, mini-swe-agent, mini-swe-agent-2-4.2
27.02 THIN
31 Claude Sonnet 4.5 Anthropic closed 38.8 47.8 13
of 32
14
claude-code, general-model, maya-v2 +3
4.26 THIN
32 Claude 4.1 Opus Anthropic closed 34.9 46.1 11
of 32
7
claude-code, mini-swe-agent, openhands +1
2.57 THIN
33 Claude 4 Opus imputed: preference Anthropic closed 30.0 49.3 10
of 32
7
diff
4.09 FULL
34 Gemini 2.5 Pro Google DeepMind closed 27.7 33.4 10
of 32
8
gemini-cli, mini-swe-agent, openhands +1
6.92 THIN
35 GPT-OSS 120B imputed: preference OpenAI unknown 24.1 41.0 8
of 32
5
diff, mini-swe-agent, terminus-2
2.25 FULL
36 Claude 3.7 Sonnet imputed: preference Anthropic closed 23.3 41.6 11
of 32
11
diff, general-model
3.18 FULL
37 Claude 4 Sonnet imputed: preference Anthropic closed 23.1 32.9 11
of 32
8
diff
6.42 FULL

Table scrolls sideways

Rule cost is how many normalised points this model would have gained had the board taken its best configuration instead of its median one. A large value is not a defect in the model and not a defect in the board: it means that model's configurations disagree with each other by a lot, so any single number for it is standing in for a wide range. Read a high rule cost as a warning that the row is less settled than its neighbours, and note that under the best-of rule those same rows would have been ranked higher on nothing but the spread of their own measurements.

Not separable marks a pair this evidence cannot order. It appears where one row has a measured preference score and the other has that category imputed, and the two sit within 1.78 index points. Section 09 gives the measurement behind that number. The band is applied one pair at a time and never chained into tie groups, because chaining is not transitive and would collapse rows tens of points apart into a single tie.

04 The strict board

The same rubric applied to the rigorous entity: one exact (model, harness, effort) tuple, with nothing rolled up. 2 of 1147 systems qualify.

#SystemHarnessEffort CompositeIntervalEvidence
1 Kimi K2.6 49.8 ±4.1 THIN
2 Claude Opus 4.6 42.4 ±2.0 THIN

What would fill it

Of 1147 systems identified, 2 hold enough evidence in all three required categories, 26 hold it in two, 245 in one, and 874 in none. 14 systems are a single measurement away from qualifying.

Every near-miss system has no harness recorded. A benchmark that publishes a harness column would create a new system rather than complete one of these, so only harness-free benchmarks appear as candidates here.

Run this benchmarkSystems it would qualify
cybench 11 of the 14 systems that are one measurement short
metr_time_horizon 11 of the 14 systems that are one measurement short
arena_agent 7 of the 14 systems that are one measurement short
apex_agents 4 of the 14 systems that are one measurement short
swe_bench_verified 2 of the 14 systems that are one measurement short
ale_bench 1 of the 14 systems that are one measurement short

These counts are alternatives, not additions. Cybench and the METR time horizon evaluation would each qualify the same 11 systems, so running both buys nothing that running either does not. Pick one. That is what makes the near-empty strict board a measurement plan rather than a dead end.

SystemOrganisationShort categoryAlready hasAny one of
Gemini 3 Flash Preview Google DeepMind agents apex_agents metr_time_horizon, cybench, arena_agent
Gemini 3 Pro Preview Google DeepMind agents apex_agents metr_time_horizon, cybench, arena_agent
Gemini 3.1 Pro Preview Google DeepMind coding ale_bench swe_bench_verified
Gemini 3.5 Flash Google agents arena_agent apex_agents, metr_time_horizon, cybench
GLM 5.1 Z.ai (Zhipu AI) agents arena_agent apex_agents, metr_time_horizon, cybench
GLM-5.2 Z.ai (Zhipu AI) agents arena_agent apex_agents, metr_time_horizon, cybench
GPT-4o (2024-11-20) OpenAI coding swe_bench_verified ale_bench
GPT-5 (2025-08-07) OpenAI agents apex_agents metr_time_horizon, cybench, arena_agent
GPT-5.1 (2025-11-13) OpenAI agents apex_agents metr_time_horizon, cybench, arena_agent
GPT-5.2 (2025-12-11) OpenAI agents apex_agents metr_time_horizon, cybench, arena_agent
Kimi K2.5 Moonshot agents apex_agents metr_time_horizon, cybench, arena_agent
Kimi K2.7 Code Moonshot coding ale_bench swe_bench_verified
Qwen3.7 Max Alibaba agents arena_agent apex_agents, metr_time_horizon, cybench
Z-Ai/GLM 5 Z.ai (Zhipu AI) agents apex_agents metr_time_horizon, cybench, arena_agent

This board is small because publishers do not report a harness consistently. Epoch's own SWE-bench Verified rows name no harness while DeepSWE's do, so those two measurements describe different configurations and are never merged. The number is published rather than hidden: a near-empty strict board is a fact about the state of public benchmark reporting, and it is the reason the model rollup exists as a second, separately labelled entity.

05 Category boards

Each classification is answered twice. First the ranking, on placement, which is the same ordering the board publishes: one snapshot gives one answer. Then every configuration measured, as evidence, so a reader can see how far a result moves when the harness or the reasoning effort changes. The second table is deliberately not a ranking.

Agents weight 35%

#ModelHead to headPlayed BenchesBest finishOverall
1 Claude Opus 5 98.6% 70 of 71 3 1 of 42 on LMArena Agent 1
2 Claude Fable 5 97.2% 69 of 71 3 1 of 47 on APEX Agents 3
3 Kimi K3 87.1% 54 of 62 2 4 of 42 on LMArena Agent 4
4 GPT-5.6 Sol 83.1% 59 of 71 3 3 of 42 on LMArena Agent 2
5 GPT-5.5 86.4% 76 of 88 3 1 of 33 on Terminal-Bench 2 10
6 Grok 4.6 74.6% 53 of 71 3 5 of 47 on APEX Agents 5
7 Claude Opus 4.7 79.5% 70 of 88 3 3 of 33 on Terminal-Bench 2 9
8 Claude Opus 4.8 71.8% 51 of 71 3 3 of 47 on APEX Agents 6
9 GPT-5.4 (2026-03-05) 86.7% 72 of 83 3 2 of 33 on Terminal-Bench 2 7
10 GLM-5.2 72.6% 45 of 62 2 9 of 47 on APEX Agents 14

Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 54 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.

Every configuration measured

This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.

367 systems with at least one qualifying benchmark: 20 FULL, 0 THIN, 347 SINGLE. 280 contributing measurements carry no measurement date. Showing 15.

clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.

A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.

RowSystemConfigurationScoreInterval BenchesWhichEvidence
1 Intelligence Indeed Agent agentic-framework / 100-steps 90.2 1 osworld SINGLE
2 Claude Mythos Preview Early as published 86.3 ±13.7 1 metr_time_horizon SINGLE
3 Claude Fable 5[1m] general-model / 100-steps 86.0 1 osworld SINGLE
4 GPT-5.5 nexau-ahe 84.7 ±4.1 1 terminalbench SINGLE
5 Pointer Agent W/ Opus 4.7 agentic-framework / 100-steps 83.6 1 osworld SINGLE
6 Claude Opus 5[1m] general-model / 100-steps 83.4 1 osworld SINGLE
7 GPT-5.5 capy 83.1 ±4.1 1 terminalbench SINGLE
8 Coasty Cua V1 general-model / 100-steps 82.8 1 osworld SINGLE
9 Holo3-35B-A3B specialized-model / 100-steps 82.6 1 osworld SINGLE
10 GPT-5.5 codex-cli 82.2 ±4.3 1 terminalbench SINGLE
11 GPT-5.5 codex 82.0 ±4.3 1 terminalbench SINGLE
12 GPT-5.4 (2026-03-05) forgecode 81.8 ±3.9 1 terminalbench SINGLE
13 Pointer Agent W/ Sonnet 4.6 agentic-framework / 100-steps 81.5 1 osworld SINGLE
14 Muse Spark 1.1 general-model / 100-steps 80.7 1 osworld SINGLE
15 Claude Opus 4.7 wozcode 80.2 ±4.1 1 terminalbench SINGLE

Coding weight 30%

#ModelHead to headPlayed BenchesBest finishOverall
1 Claude Opus 5 97.8% 91 of 93 3 1 of 23 on DeepSWE 1
2 GPT-5.6 Sol 95.7% 89 of 93 3 1 of 82 on ALE-Bench 2
3 Claude Fable 5 94.6% 88 of 93 3 1 of 25 on FrontierCode 3
4 Kimi K3 82.8% 77 of 93 3 8 of 82 on ALE-Bench 4
5 Grok 4.6 81.7% 76 of 93 3 3 of 25 on FrontierCode 5
6 GPT-5.5 78.5% 73 of 93 3 7 of 82 on ALE-Bench 10
7 GPT-5.6 Terra 76.9% 71.5 of 93 3 4 of 82 on ALE-Bench 12
8 GLM-5.3 75.7% 53 of 70 2 3 of 23 on DeepSWE 8
9 Claude Opus 4.8 72.6% 67.5 of 93 3 12 of 82 on ALE-Bench 6
10 GPT-5.6 Luna 68.8% 64 of 93 3 5 of 82 on ALE-Bench 16

Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 54 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.

Every configuration measured

This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.

275 systems with at least one qualifying benchmark: 16 FULL, 1 THIN, 258 SINGLE. 212 contributing measurements carry no measurement date. Showing 15.

clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.

A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.

RowSystemConfigurationScoreInterval BenchesWhichEvidence
1 GPT-5 (2025-08-07) diff / high 88.0 1 aider_polyglot SINGLE
2 GPT-5 (2025-08-07) diff / medium 86.7 1 aider_polyglot SINGLE
3 OpenAI o3-pro (2025-06-10) diff / high 84.9 1 aider_polyglot SINGLE
4 Claude Opus 4.7 max 83.5 ±3.3 1 swe_bench_verified SINGLE
5 Gemini 2.5 Pro Preview 0605 diff-fenced / 32k 83.1 1 aider_polyglot SINGLE
6 GPT-5 (2025-08-07) diff / low 81.3 1 aider_polyglot SINGLE
7 o3 (2025-04-16) diff / high 81.3 1 aider_polyglot SINGLE
8 GPT-5.5 Pre Release xhigh 80.6 ±3.5 1 swe_bench_verified SINGLE
9 Grok 4.0709 diff 79.6 1 aider_polyglot SINGLE
10 Grok 4.0709 diff / high 79.6 1 aider_polyglot SINGLE
11 Gemini 2.5 Pro Preview 0605 diff-fenced 79.1 1 aider_polyglot SINGLE
12 DeepSeek V4 Pro max 77.6 ±3.7 1 swe_bench_verified SINGLE
13 Gemini 2.5 Pro Preview 0506 diff-fenced 76.9 1 aider_polyglot SINGLE
14 o3 (2025-04-16) diff / medium 76.9 1 aider_polyglot SINGLE
15 o3 (2025-04-16) diff 76.9 1 aider_polyglot SINGLE

Reasoning weight 25%

#ModelHead to headPlayed BenchesBest finishOverall
1 GPT-5.5 Pre Release 99.1% 382.5 of 386 3 1 of 67 on FrontierMath null
2 GPT-5.5 Pro Pre Release 98.6% 380.5 of 386 3 2 of 161 on OTIS Mock AIME 2024-2025 null
3 GPT-5.4 Pro (2026-03-05) 97.7% 317.5 of 325 4 2 of 163 on GPQA Diamond (Epoch's own run) null
4 Claude Opus 5 96.5% 371.5 of 385 3 1 of 66 on ARC-AGI-2 1
5 Qwen3.8 Max 96.6% 311 of 322 2 4 of 161 on OTIS Mock AIME 2024-2025 null
6 Gemini 3.7 Flash 94.8% 365 of 385 3 1 of 163 on GPQA Diamond (Epoch's own run) null
7 Grok 4.6 93.6% 360.5 of 385 3 7 of 163 on GPQA Diamond (Epoch's own run) 5
8 Gemini 3.1 Pro Preview 93.5% 453.5 of 485 5 1 of 37 on Humanity's Last Exam 17
9 Grok 4.5 92.9% 357.5 of 385 3 6 of 163 on GPQA Diamond (Epoch's own run) 11
10 DeepSeek V4 Pro 0813 92.5% 356 of 385 3 5 of 161 on OTIS Mock AIME 2024-2025 null

Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 168 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.

Every configuration measured

This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.

398 systems with at least one qualifying benchmark: 265 FULL, 1 THIN, 132 SINGLE. 274 contributing measurements carry no measurement date. Showing 15.

clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.

A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.

RowSystemConfigurationScoreInterval BenchesWhichEvidence
1 Qwen3.8 Max xhigh 96.1 ±1.7 2 gpqa_diamond, otis_mock_aime FULL
2 Claude Opus 5 as published 95.4 ±2.8 2 gpqa_diamond, otis_mock_aime FULL
3 GPT-5.6 Sol max 95.3 ±1.0 3 arc_agi_2, gpqa_diamond, otis_mock_aime FULL
4 Gemini 3.1 Pro Preview high 95.0 ±3.4 2 gpqa_diamond, otis_mock_aime FULL
5 Claude Opus 5 max 94.0 ±1.3 3 arc_agi_2, gpqa_diamond, otis_mock_aime FULL
6 DeepSeek V4 Pro high 93.2 ±3.6 2 gpqa_diamond, otis_mock_aime FULL
7 Qwen3.7 Max as published 93.2 ±3.6 2 gpqa_diamond, otis_mock_aime FULL
8 DeepSeek V4 Pro max 93.2 ±2.6 2 gpqa_diamond, otis_mock_aime FULL
9 Claude Sonnet 5 xhigh 92.6 ±2.8 2 gpqa_diamond, otis_mock_aime FULL
10 Gemini 3 Flash Preview high 92.5 ±3.7 2 gpqa_diamond, otis_mock_aime FULL
11 GPT-5.6 Terra max 92.3 ±1.0 3 arc_agi_2, gpqa_diamond, otis_mock_aime FULL
12 Gemini 3.7 Flash high 92.2 ±1.5 3 arc_agi_2, gpqa_diamond, otis_mock_aime FULL
13 GLM-5.3-Flash max 92.0 ±2.7 2 gpqa_diamond, otis_mock_aime FULL
14 Kimi K2.7 Code as published 91.7 ±3.8 2 gpqa_diamond, otis_mock_aime FULL
15 Claude Fable 5 max 91.6 ±1.6 3 arc_agi_2, gpqa_diamond, otis_mock_aime FULL

Preference weight 10%

#ModelHead to headPlayed BenchesBest finishOverall
1 Claude Fable 5 97.4% 187 of 192 2 1 of 158 on LMArena Text (style-controlled) 3
2 Claude Opus 5 High 95.3% 183 of 192 2 4 of 100 on LMArena WebDev null
3 Claude Opus 5 Max 94.8% 182 of 192 2 1 of 100 on LMArena WebDev null
4 Kimi K3 Max 94.8% 182 of 192 2 2 of 100 on LMArena WebDev null
5 Claude Opus 4.7 High 91.7% 176 of 192 2 3 of 158 on LMArena Text (style-controlled) null
6 Claude Opus 4.6 High 91.1% 175 of 192 2 2 of 158 on LMArena Text (style-controlled) null
7 Gemini 3.7 Flash High 91.1% 175 of 192 2 9 of 158 on LMArena Text (style-controlled) null
8 Claude Opus 4.7 90.6% 174 of 192 2 6 of 158 on LMArena Text (style-controlled) 9
9 GLM-5.3 Max 89.6% 172 of 192 2 8 of 100 on LMArena WebDev null
10 Qwen3.8 Max 89.6% 172 of 192 2 3 of 100 on LMArena WebDev null

Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 97 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.

Every configuration measured

This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.

419 systems with at least one qualifying benchmark: 0 FULL, 93 THIN, 326 SINGLE. 512 contributing measurements carry no measurement date. Showing 15.

clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.

A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.

RowSystemConfigurationScoreInterval BenchesWhichEvidence
1 Claude Opus 5 Max as published 85.2 ±1.2 2 arena_text_style_control, arena_webdev
sources disagree
THIN
2 Kimi K3 Max as published 84.2 ±1.4 2 arena_text_style_control, arena_webdev THIN
3 Muse Spark as published 84.1 ±0.8 1 arena_text_style_control SINGLE
4 Claude Opus 5 High as published 83.7 ±1.0 2 arena_text_style_control, arena_webdev THIN
5 GPT-5.5 High as published 83.2 ±0.6 1 arena_text_style_control SINGLE
6 Qwen3.8 Max as published 83.2 ±1.6 2 arena_text_style_control, arena_webdev THIN
7 GPT-5.6 Sol Xhigh as published 83.2 ±0.7 1 arena_text_style_control SINGLE
8 Claude Fable 5 as published 82.5 ±4.3 2 arena_text_style_control, arena_webdev
sources disagree
THIN
9 GPT-5.5 as published 82.4 ±0.6 1 arena_text_style_control SINGLE
10 GPT-5.4 High as published 82.4 ±0.6 1 arena_text_style_control SINGLE
11 GPT-5.2 Chat Latest (2026-02-10) as published 82.3 ±0.6 1 arena_text_style_control SINGLE
12 Grok 4.20 Beta1 as published 82.1 ±0.7 1 arena_text_style_control SINGLE
13 Qwen3.7 Max Preview as published 82.0 ±1.4 1 arena_text_style_control SINGLE
14 GPT-5.5 Instant as published 81.9 ±0.7 1 arena_text_style_control SINGLE
15 Grok 4.20 Multi Agent Beta 0309 as published 81.5 ±0.5 1 arena_text_style_control SINGLE

06 Every benchmark behind the numbers

Who measured it, who moved it, on what terms, on what scale, and how many of its rows survived to enter a score.

IDBenchmarkCategoryEvaluatorTransport LicenceMetricAnchorsObservedInterval AttributionContaminationIngestedScored
terminalbench Terminal-Bench 2 Agents Terminal-Bench Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.03 to 0.85 stderr row low 212 104
apex_agents APEX Agents Agents APEX Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.01 to 0.45 stderr file low 63 62
metr_time_horizon METR time horizons Agents METR Epoch AI Benchmarking Hub CC-BY-4.0 time_horizon_min 0.00 to 3.50 0.05 to 1044.78 ci row low 50 36
cybench Cybench Agents Cybench Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.05 to 0.93 none row medium 22 16
arena_agent LMArena Agent Agents LMArena LMArena leaderboard dataset CC-BY-4.0 ips_score -0.05 to 0.25 -0.21 to 0.13 ci own n/a 54 54
swe_bench_verified SWE-bench Verified (Epoch's own run) Coding Epoch AI Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.31 to 0.83 stderr own medium 41 40
aider_polyglot Aider Polyglot Coding Aider Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.04 to 0.88 none row medium 72 72
deepswe DeepSWE Coding Datacurve Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.02 to 0.74 ci_halfwidth row medium 62 61
frontiercode FrontierCode Coding Cognition Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.08 to 0.53 none row low 28 28
gpqa_diamond GPQA Diamond (Epoch's own run) Reasoning Epoch AI Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.13 to 0.95 stderr own high 284 280
frontiermath FrontierMath Reasoning Epoch AI Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.00 to 0.52 stderr own low 101 101
hle Humanity's Last Exam Reasoning Humanity’s Last Exam (CAIS / Scale AI) Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.03 to 0.46 stderr file low 59 53
otis_mock_aime OTIS Mock AIME 2024-2025 Reasoning Epoch AI Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.00 to 1.00 stderr own medium 256 256
arc_agi_2 ARC-AGI-2 Reasoning ARC Prize Epoch AI Benchmarking Hub CC-BY-4.0 pass_rate 0.00 to 1.00 0.00 to 0.93 none file low 216 216
arena_text_style_control LMArena Text (style-controlled) Preference LMArena LMArena leaderboard dataset CC-BY-4.0 bt_rating 900.00 to 1600.00 952.54 to 1507.48 ci own n/a 395 395
arena_webdev LMArena WebDev Preference LMArena LMArena leaderboard dataset CC-BY-4.0 bt_rating 1000.00 to 1800.00 1080.11 to 1690.64 ci own n/a 117 117
ale_bench ALE-Bench Coding ALE-Bench (Sakana AI with AtCoder) Epoch AI Benchmarking Hub CC-BY-4.0 perf_rating 0.00 to 3500.00 137.78 to 2176.88 none file low 110 110
terminalbench_4 Terminal-Bench 4.0 Agents Terminal-Bench Terminal-Bench 4.0 leaderboard submissions Apache-2.0 pass_rate 0.00 to 1.00 ci_halfwidth own low 10 10
tau2_airline Tau2-Bench Airline Agents Tau2-Bench (Sierra) Tau2-Bench leaderboard MIT pass_rate 0.00 to 1.00 none own low 0 0
tau2_retail Tau2-Bench Retail Agents Tau2-Bench (Sierra) Tau2-Bench leaderboard MIT pass_rate 0.00 to 1.00 none own low 0 0
tau2_telecom Tau2-Bench Telecom Agents Tau2-Bench (Sierra) Tau2-Bench leaderboard MIT pass_rate 0.00 to 1.00 none own low 0 0
tau2_banking Tau2-Bench Banking Agents Tau2-Bench (Sierra) Tau2-Bench leaderboard MIT pass_rate 0.00 to 1.00 none own low 0 0
osworld OSWorld Agents OSWorld (XLANG Lab) OSWorld leaderboard none stated pass_rate 0.00 to 1.00 none own low 164 110
browsecomp BrowseComp Agents various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 7 1
mcp_atlas MCP-Atlas Agents various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 4 0
tool_decathlon Tool-Decathlon Agents various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 2 0
cybergym CyberGym Agents various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 0 0
swe_bench_pro SWE-bench Pro Coding various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 2 0
swe_bench_multilingual SWE-bench Multilingual Coding various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 1 1
aime AIME Reasoning various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 3 0
hmmt HMMT Reasoning various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 4 0
imo_answerbench IMOAnswerBench Reasoning various; recorded per row Model benchmark dossiers facts, not expression pass_rate 0.00 to 1.00 none row low 4 2

Table scrolls sideways

Ingested and Scored are deliberately different numbers, and so is the row count in the source signature printed in section 11. A source file has some number of rows; some of those rows have no usable value or no identifiable model and never become an observation; and some observations are then refused entry to a score by the rights gate or the provenance gate. Three numbers, narrowing at each step. Section 07 accounts for the difference. Contamination is our recorded judgement of how likely a benchmark's tasks are to have appeared in training data. It is a risk flag, not a measurement, and a high rating is a reason to weigh that benchmark's contribution more carefully rather than to discard it.

Evaluator is who ran the benchmark. Transport is who we obtained it from. They are recorded separately and both are credited, because crediting only the aggregator would misstate who did the work. Attribution says how the evaluator is known: own means the transport ran the evaluation itself, row means the file carries a per-row source column so each row can be screened for vendor self-reports, and file means the file has no attribution column at all, so the evaluator is a reviewed assertion about the whole file and individual rows cannot be screened. Rows from a file-level source are counted and reported in the next section.

Correlation families

Benchmarks that measure the same thing are averaged inside their family before the category mean, so four variants of one task distribution count as one piece of evidence rather than four.

07 What was refused, and what is uncertain

An exclusion the reader cannot see is the same as no exclusion at all.

Refused entry to the score

StateRowsWhy
UNATTRIBUTED114 The file carries a source column and this row left it blank, so nobody can be named as the evaluator.
CLAIMED104 A vendor self-report. Stored and displayable, never scored.
rights0 The licence on the source does not permit us to restate its numbers. We may read it and link to it, and we do, but it never enters a score.

Known weaknesses in what did get scored

MeasureCountWhat it means
Unscreened rows439 From files with no attribution column. The evaluator is named for the file as a whole, but no individual row can be checked against it.
Unmapped identifiers836 Model identifiers seen in the data that no reviewed alias covers. They are recorded rather than guessed at, because a wrong merge is worse than a missing row.
Metadata conflicts740 Model keys where providers disagree about price, context window or open weights. Where they disagree we publish nothing for that field rather than picking one.
Systems seen but not ranked1145 Identified, scored where possible, and short of the entry requirement in at least one required category.

08 Method

Weights
Agents 35%, Coding 30%, Reasoning 25%, Preference 10%. knowledge and long_context carry zero weight in this rubric version: they are ingested and displayed, never scored.
Required categories
agents, coding, reasoning. A model missing any of them has no index. preference is optional, and its absence is declared on the row.
Entry requirement
2 qualifying benchmarks in each required category, from at least 2 distinct correlation families before a category counts as FULL rather than THIN.
Normalisation
Every raw value is mapped onto 0 to 100 against fixed anchors, never against the current field. Normalising against the live field would move every published score whenever an unrelated model launched. Anchors are per benchmark wherever the scale differs, because two rating systems do not share a scale.
Recency
Measured on the date of MEASUREMENT, not the model's release date, with a default cutoff of 18 months. Filtering on release date would delete every comparison baseline from the board.
Uncertainty
Published intervals are used as published and never recomputed. Where members of a correlation family disagree, the band is widened by half their spread rather than being allowed to shrink, so disagreement between two measurements of the same thing shows up as less confidence instead of vanishing into a mean. A published interval of width zero is treated as saturation evidence, not precision evidence.
Ranking
weighted mean of each category's head-to-head position within its published field. Positions, not scores: a category counts for where a model finished against everyone who sat the same benchmarks. The composite the previous rule produced is kept on every row and printed beside the index.. Confidence-interval overlap is not used to group rows into tiers, because overlap is not transitive: a row can overlap its neighbour, and that neighbour overlap its own, while the two ends are clearly apart.
Rollup
median observed configuration per benchmark (lower median, always a real measurement).
Self-reports
A value whose source is a vendor technical report, system card, model card or launch announcement is recorded as CLAIMED and never enters a score.

09 What this report cannot tell you

Every limit below is measured, not estimated, and every one of them is a reason to read a specific number with more care.

The rubric is not calibrated

UNCALIBRATED - anchors and thresholds are placeholders until the 90-day backfill study runs. Every published number must carry this label. Anchors and thresholds are placeholders. Ordering within the board is far more trustworthy than the absolute value of any index.

Benchmarks of different difficulty are averaged together

A category score is the mean across its benchmark families, and those families are not equally hard. In Coding alone this report scores SWE-bench Verified (Epoch's own run), Aider Polyglot, DeepSWE, FrontierCode, ALE-Bench, SWE-bench Pro, SWE-bench Multilingual, whose normalised distributions sit at very different heights. Until the calibration study equalises them, being measured on a harder benchmark costs a model points that a reader would naturally attribute to weaker ability.

The missing preference category is biased, not merely absent

RE-MEASURED on 2026-08-29 when the index moved from composite category scores to head-to-head positions, because the previous band of 2.17 was derived in composite points and could not be carried onto a different scale. The measurement is the same one, run again: for the 13 ranked models holding both a qualifying preference position and all three required categories, compare the measured preference position with the position implied by the weighted mean of the others. THE BIAS DISAPPEARED. It was +21.66 composite points, 8 of 8 in the same direction. It is now -1.64 position points with an sd of 17.79 and only 5 of 13 positive, which is indistinguishable from zero. That is exactly what the previous version of this note predicted would happen: it argued the offset was an artefact of incommensurable anchors rather than a property of models, since preference is a Bradley-Terry rating normalised against placeholder anchors of its own while the other categories are pass rates against 0 to 1. Positions are scale-free, so the artefact had nothing left to live in. The hypothesis was recorded before the test and the test agreed with it. The band therefore now describes NOISE, not bias. Imputing preference moves a row by 17.79 position points one standard deviation, which at the 10% preference weight is 1.78 index points, so an ordering finer than that between an imputed row and a measured row is not something this board can assert. On the live snapshot the band contains 18 of 561 pairs, 10 of them with exactly one side imputed, against a median adjacent gap of 1.95 index points.

There is no longer an offset to correct, and that is the finding rather than a convenience. The previous version of this field declined to add back a +21.66 point offset on the grounds that it was an artefact of incommensurable anchors rather than a property of models: preference is a Bradley-Terry rating normalised against placeholder anchors of its own, 900 to 1600 for the text arena and 1000 to 1800 for the web development arena, which run on different scales and so cannot share one, while the other categories are pass rates against 0 to 1. Moving the index to positions removed the anchors from the comparison, and the offset went with them. The band is now a width and not a shift, so there is nothing to add back and adding anything would be fitting noise.

Measured on 2026-08-29 over 13 models. That is a thin sample, and it is re-measured as coverage grows.

Some rows cannot be screened

439 scored rows come from files that carry no attribution column. The evaluator is named for the file, and the file is reviewed, but no individual row can be checked. They are marked rather than dropped, because dropping them would remove real coverage and hide the fact that the source does not publish attribution.

Some measurements carry no date

1278 contributing measurements across the category boards were published without a measurement date. They are counted and shown rather than assumed current or silently dropped.

Absence of a model is not a judgement about it

A model can be missing because no one has published two independent measurements of it in a required category. That is a statement about public benchmark coverage, not about the model.

10 Corrections

Every figure here traces to a snapshot id and a source manifest id, both printed at the top of this report and repeated in the footer. Quote them when reporting a problem and the exact inputs can be replayed to the byte. The same snapshot is available as a machine-readable document, so any figure on this page can be checked against the object it was rendered from without taking our word for it.

Corrections go to ferroxlabs.com. A correction that changes a published number is described when it is applied, not quietly folded into the next snapshot.

11 Attribution and licences

Required on any page that shows these numbers.

Sources in this snapshot

SourceStatusObservationsSignature
arena ok 566 arena_agent:54,arena_text_style_control:395,arena_webdev:117
dossier ok 61 dossiers:3,kept:61,dropped:178,hash:c9db018492921a02
epoch ok 1546 aider_polyglot:77,ale_bench:110,apex_agents:62,arc_agi_2:216,cybench:22,deepswe:61,frontiercode:33,frontiermath:101,gpqa_diamond:280,hle:51,metr_time_horizon:50,otis_mock_aime:256,swe_bench_verified:35,terminalbench:204
modelsdev ok 0 providers:211,models:7488
osworld ok 160 verified:1076,self:56,kept:160,dropped:972,hash:65022c7f3262daac
tbench4 ok 10 submissions:10,kept:10,dropped:0,tree:de3608c2dd501f81