THE FORGE / RANKINGS
This board ranks public measurements, not models. Every number names the benchmark that produced it, the evaluator who ran that benchmark, the source it came through and the licence it was published under. Vendor self-reports are recorded and never scored. Ferrox has run none of these evaluations itself; where a measurement cannot be attributed to a named evaluator, the row says so and does not enter a score.
Anchors and thresholds are placeholders until the 90-day backfill study runs. Every published number must carry this label.
Read the ordering, not the absolute value. The rubric is uncalibrated, so the gap between two rows is far more trustworthy than the number printed on either of them.
2343
Observations ingested
every row read from every source
2125
Entered a score
90.7% of what was read
1147
Systems identified
model, harness and effort tuples
37
Models ranked
of 610 models seen
2
Systems ranked
under the strict tuple entity
32
Benchmarks in the score
each one declared in the stat contract
A system is one exact combination of model, harness and effort. It is the rigorous unit: two measurements are joined only when all three match, so a run whose harness was never published is never merged with one whose harness was.
A model is that model across every configuration measured. Rolling up to it is a declared operation with a stated rule, not a guess, and every model row names how many configurations sit behind it and which harnesses they used.
The rollup rule is stated, not implied: median observed configuration per benchmark (lower median, always a real measurement). Taking the best configuration instead would reward a model for having been measured more often, because the maximum of more draws is higher whether or not the model is better.
A Ferrox Index score requires at least 2 qualifying benchmarks in each of agents, coding, reasoning. Systems below that bar appear on the category boards with their evidence count shown, and have no overall score.
The interval printed beside an index is built only from the benchmarks that published one. Five of the benchmarks scored here publish no uncertainty at all, so for most rows the interval is a lower bound: the true uncertainty is at least this wide and can only be wider. Rows in that state are marked "at least". Publishing no interval at all would have thrown away the real information the other sources did provide.
Rows are ordered by the point estimate. The interval is printed beside it and is never used as the sort key. Sorting by the top of the interval instead would make a wider interval a rank bonus, and would reward a source for publishing a vaguer number.
Models, rolled up across every configuration measured. 37 of 610 models clear the entry requirement.
| # | Model | Organisation | Weights | Ferrox Index | Old composite | Benches | Configs | Rule cost | Evidence |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 imputed: preference | Anthropic | closed | 99.5 | 66.7 | 9 of 32 |
10 claude-code, mini-swe-agent |
1.17 | FULL |
| 2 | GPT-5.6 Sol imputed: preference not separable from 3 | OpenAI | closed | 95.5 | 61.3 | 9 of 32 |
13 codex, mini-swe-agent |
4.12 | FULL |
| 3 | Claude Fable 5 not separable from 2, 4 | Anthropic | closed | 94.8 | 65.8 | 11 of 32 |
13 claude-code, mini-swe-agent |
1.45 | THIN |
| 4 | Kimi K3 imputed: preference not separable from 3 | Moonshot | open | 93.7 | 56.5 | 8 of 32 |
6 mini-swe-agent |
3.50 | FULL |
| 5 | Grok 4.6 imputed: preference | xAI | closed | 92.8 | 53.5 | 9 of 32 |
11 grok-build, mini-swe-agent |
2.58 | FULL |
| 6 | Claude Opus 4.8 not separable from 7 | Anthropic | closed | 87.1 | 51.3 | 12 of 32 |
12 claude-code, mini-swe-agent |
6.63 | THIN |
| 7 | GPT-5.4 (2026-03-05) imputed: preference not separable from 6, 9 | OpenAI | unknown | 85.5 | 60.8 | 11 of 32 |
8 forgecode, mini-swe-agent |
3.38 | FULL |
| 8 | GLM-5.3 imputed: preference not separable from 9 | Z.ai (Zhipu AI) | unknown | 84.9 | 57.0 | 6 of 32 |
4 claude-code, mini-swe-agent |
0.00 | FULL |
| 9 | Claude Opus 4.7 not separable from 7, 8 | Anthropic | closed | 84.0 | 57.6 | 13 of 32 |
8 claude-code, wozcode |
5.75 | THIN |
| 10 | GPT-5.5 not separable from 12 | OpenAI | closed | 81.6 | 58.0 | 11 of 32 |
17 capy, clnkr, codex +3 |
8.98 | THIN |
| 11 | Grok 4.5 not separable from 12 | xAI | closed | 81.0 | 51.1 | 11 of 32 |
7 grok-build, mini-swe-agent |
0.00 | THIN |
| 12 | GPT-5.6 Terra imputed: preference not separable from 10, 11 | OpenAI | closed | 80.1 | 48.7 | 8 of 32 |
12 codex, mini-swe-agent |
12.36 | FULL |
| 13 | Claude Sonnet 5 imputed: preference | Anthropic | closed | 74.2 | 48.3 | 8 of 32 |
11 claude-code, mini-swe-agent |
10.19 | FULL |
| 14 | GLM-5.2 imputed: preference not separable from 15, 17 | Z.ai (Zhipu AI) | unknown | 71.1 | 43.6 | 9 of 32 |
8 mini-swe-agent |
5.84 | FULL |
| 15 | Claude Opus 4.6 not separable from 14, 16 | Anthropic | closed | 70.8 | 52.3 | 13 of 32 |
11 capy, claude-code, droid +3 |
3.81 | THIN |
| 16 | GPT-5.6 Luna imputed: preference not separable from 15, 17 | OpenAI | closed | 70.8 | 40.0 | 8 of 32 |
12 codex, mini-swe-agent |
23.53 | FULL |
| 17 | Gemini 3.1 Pro Preview not separable from 14, 16 | Google DeepMind | closed | 69.9 | 48.4 | 13 of 32 |
7 forge-code, forgecode, gemini-cli +2 |
0.55 | THIN |
| 18 | Kimi K2.6 not separable from 19, 20, 21 | Moonshot | unknown | 65.9 | 56.1 | 10 of 32 |
2 general-model |
0.00 | THIN |
| 19 | GPT-5.2 (2025-12-11) imputed: preference not separable from 18 | OpenAI | unknown | 65.5 | 50.5 | 9 of 32 |
8 codex-cli, droid |
4.82 | FULL |
| 20 | Gemini 3 Pro Preview imputed: preference not separable from 18 | Google DeepMind | closed | 65.3 | 51.4 | 9 of 32 |
5 codebrain-1-5, droid, ii-agent +1 |
1.10 | FULL |
| 21 | Gemini 3 Flash Preview imputed: preference not separable from 18 | Google DeepMind | closed | 65.0 | 47.6 | 8 of 32 |
4 gemini-cli, junie-cli |
13.14 | FULL |
| 22 | Claude Sonnet 4.6 not separable from 23 | Anthropic | closed | 62.6 | 50.2 | 14 of 32 |
10 claude-code, general-model, mini-swe-agent +1 |
5.46 | THIN |
| 23 | Kimi K2.7 Code imputed: preference not separable from 22, 24 | Moonshot | open | 61.5 | 44.7 | 8 of 32 |
3 mini-swe-agent |
0.00 | FULL |
| 24 | Claude 4.5 Opus not separable from 23 | Anthropic | closed | 61.4 | 50.7 | 12 of 32 |
7 droid, letta-code |
7.30 | THIN |
| 25 | GLM 5.1 | Z.ai (Zhipu AI) | unknown | 58.7 | 53.3 | 13 of 32 |
7 ascend-a3-dual-node, ascend-a5-w4a4c8, gpu-baseline;-10-groups-over-3-rounds +3 |
10.17 | THIN |
| 26 | GPT-5.1 (2025-11-13) imputed: preference | OpenAI | closed | 47.4 | 38.7 | 9 of 32 |
6 terminus-2 |
9.46 | FULL |
| 27 | GPT-5 (2025-08-07) imputed: preference | OpenAI | unknown | 45.2 | 46.8 | 11 of 32 |
12 codex-cli, diff, mini-swe-agent +2 |
4.60 | FULL |
| 28 | Kimi K2.5 imputed: preference | Moonshot | unknown | 43.4 | 36.3 | 6 of 32 |
2 general-model |
0.00 | FULL |
| 29 | o3 (2025-04-16) imputed: preference not separable from 30 | OpenAI | unknown | 42.2 | 43.3 | 11 of 32 |
7 diff |
2.07 | FULL |
| 30 | MiniMax-M3 not separable from 29 | MiniMax | open | 40.6 | 45.6 | 9 of 32 |
5 general-model, mini-swe-agent, mini-swe-agent-2-4.2 |
27.02 | THIN |
| 31 | Claude Sonnet 4.5 | Anthropic | closed | 38.8 | 47.8 | 13 of 32 |
14 claude-code, general-model, maya-v2 +3 |
4.26 | THIN |
| 32 | Claude 4.1 Opus | Anthropic | closed | 34.9 | 46.1 | 11 of 32 |
7 claude-code, mini-swe-agent, openhands +1 |
2.57 | THIN |
| 33 | Claude 4 Opus imputed: preference | Anthropic | closed | 30.0 | 49.3 | 10 of 32 |
7 diff |
4.09 | FULL |
| 34 | Gemini 2.5 Pro | Google DeepMind | closed | 27.7 | 33.4 | 10 of 32 |
8 gemini-cli, mini-swe-agent, openhands +1 |
6.92 | THIN |
| 35 | GPT-OSS 120B imputed: preference | OpenAI | unknown | 24.1 | 41.0 | 8 of 32 |
5 diff, mini-swe-agent, terminus-2 |
2.25 | FULL |
| 36 | Claude 3.7 Sonnet imputed: preference | Anthropic | closed | 23.3 | 41.6 | 11 of 32 |
11 diff, general-model |
3.18 | FULL |
| 37 | Claude 4 Sonnet imputed: preference | Anthropic | closed | 23.1 | 32.9 | 11 of 32 |
8 diff |
6.42 | FULL |
Table scrolls sideways
Rule cost is how many normalised points this model would have gained had the board taken its best configuration instead of its median one. A large value is not a defect in the model and not a defect in the board: it means that model's configurations disagree with each other by a lot, so any single number for it is standing in for a wide range. Read a high rule cost as a warning that the row is less settled than its neighbours, and note that under the best-of rule those same rows would have been ranked higher on nothing but the spread of their own measurements.
Not separable marks a pair this evidence cannot order. It appears where one row has a measured preference score and the other has that category imputed, and the two sit within 1.78 index points. Section 09 gives the measurement behind that number. The band is applied one pair at a time and never chained into tie groups, because chaining is not transitive and would collapse rows tens of points apart into a single tie.
The same rubric applied to the rigorous entity: one exact (model, harness, effort) tuple, with nothing rolled up. 2 of 1147 systems qualify.
| # | System | Harness | Effort | Composite | Interval | Evidence |
|---|---|---|---|---|---|---|
| 1 | Kimi K2.6 | — | — | 49.8 | ±4.1 | THIN |
| 2 | Claude Opus 4.6 | — | — | 42.4 | ±2.0 | THIN |
Of 1147 systems identified, 2 hold enough evidence in all three required categories, 26 hold it in two, 245 in one, and 874 in none. 14 systems are a single measurement away from qualifying.
Every near-miss system has no harness recorded. A benchmark that publishes a harness column would create a new system rather than complete one of these, so only harness-free benchmarks appear as candidates here.
| Run this benchmark | Systems it would qualify | |
|---|---|---|
| cybench | 11 | of the 14 systems that are one measurement short |
| metr_time_horizon | 11 | of the 14 systems that are one measurement short |
| arena_agent | 7 | of the 14 systems that are one measurement short |
| apex_agents | 4 | of the 14 systems that are one measurement short |
| swe_bench_verified | 2 | of the 14 systems that are one measurement short |
| ale_bench | 1 | of the 14 systems that are one measurement short |
These counts are alternatives, not additions. Cybench and the METR time horizon evaluation would each qualify the same 11 systems, so running both buys nothing that running either does not. Pick one. That is what makes the near-empty strict board a measurement plan rather than a dead end.
| System | Organisation | Short category | Already has | Any one of |
|---|---|---|---|---|
| Gemini 3 Flash Preview | Google DeepMind | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| Gemini 3 Pro Preview | Google DeepMind | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| Gemini 3.1 Pro Preview | Google DeepMind | coding | ale_bench | swe_bench_verified |
| Gemini 3.5 Flash | agents | arena_agent | apex_agents, metr_time_horizon, cybench | |
| GLM 5.1 | Z.ai (Zhipu AI) | agents | arena_agent | apex_agents, metr_time_horizon, cybench |
| GLM-5.2 | Z.ai (Zhipu AI) | agents | arena_agent | apex_agents, metr_time_horizon, cybench |
| GPT-4o (2024-11-20) | OpenAI | coding | swe_bench_verified | ale_bench |
| GPT-5 (2025-08-07) | OpenAI | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| GPT-5.1 (2025-11-13) | OpenAI | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| GPT-5.2 (2025-12-11) | OpenAI | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| Kimi K2.5 | Moonshot | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
| Kimi K2.7 Code | Moonshot | coding | ale_bench | swe_bench_verified |
| Qwen3.7 Max | Alibaba | agents | arena_agent | apex_agents, metr_time_horizon, cybench |
| Z-Ai/GLM 5 | Z.ai (Zhipu AI) | agents | apex_agents | metr_time_horizon, cybench, arena_agent |
This board is small because publishers do not report a harness consistently. Epoch's own SWE-bench Verified rows name no harness while DeepSWE's do, so those two measurements describe different configurations and are never merged. The number is published rather than hidden: a near-empty strict board is a fact about the state of public benchmark reporting, and it is the reason the model rollup exists as a second, separately labelled entity.
Each classification is answered twice. First the ranking, on placement, which is the same ordering the board publishes: one snapshot gives one answer. Then every configuration measured, as evidence, so a reader can see how far a result moves when the harness or the reasoning effort changes. The second table is deliberately not a ranking.
| # | Model | Head to head | Played | Benches | Best finish | Overall |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 98.6% | 70 of 71 | 3 | 1 of 42 on LMArena Agent | 1 |
| 2 | Claude Fable 5 | 97.2% | 69 of 71 | 3 | 1 of 47 on APEX Agents | 3 |
| 3 | Kimi K3 | 87.1% | 54 of 62 | 2 | 4 of 42 on LMArena Agent | 4 |
| 4 | GPT-5.6 Sol | 83.1% | 59 of 71 | 3 | 3 of 42 on LMArena Agent | 2 |
| 5 | GPT-5.5 | 86.4% | 76 of 88 | 3 | 1 of 33 on Terminal-Bench 2 | 10 |
| 6 | Grok 4.6 | 74.6% | 53 of 71 | 3 | 5 of 47 on APEX Agents | 5 |
| 7 | Claude Opus 4.7 | 79.5% | 70 of 88 | 3 | 3 of 33 on Terminal-Bench 2 | 9 |
| 8 | Claude Opus 4.8 | 71.8% | 51 of 71 | 3 | 3 of 47 on APEX Agents | 6 |
| 9 | GPT-5.4 (2026-03-05) | 86.7% | 72 of 83 | 3 | 2 of 33 on Terminal-Bench 2 | 7 |
| 10 | GLM-5.2 | 72.6% | 45 of 62 | 2 | 9 of 47 on APEX Agents | 14 |
Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 54 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.
This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.
367 systems with at least one qualifying benchmark: 20 FULL, 0 THIN, 347 SINGLE. 280 contributing measurements carry no measurement date. Showing 15.
clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.
A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.
| Row | System | Configuration | Score | Interval | Benches | Which | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Intelligence Indeed Agent | agentic-framework / 100-steps | 90.2 | — | 1 | osworld | SINGLE |
| 2 | Claude Mythos Preview Early | as published | 86.3 | ±13.7 | 1 | metr_time_horizon | SINGLE |
| 3 | Claude Fable 5[1m] | general-model / 100-steps | 86.0 | — | 1 | osworld | SINGLE |
| 4 | GPT-5.5 | nexau-ahe | 84.7 | ±4.1 | 1 | terminalbench | SINGLE |
| 5 | Pointer Agent W/ Opus 4.7 | agentic-framework / 100-steps | 83.6 | — | 1 | osworld | SINGLE |
| 6 | Claude Opus 5[1m] | general-model / 100-steps | 83.4 | — | 1 | osworld | SINGLE |
| 7 | GPT-5.5 | capy | 83.1 | ±4.1 | 1 | terminalbench | SINGLE |
| 8 | Coasty Cua V1 | general-model / 100-steps | 82.8 | — | 1 | osworld | SINGLE |
| 9 | Holo3-35B-A3B | specialized-model / 100-steps | 82.6 | — | 1 | osworld | SINGLE |
| 10 | GPT-5.5 | codex-cli | 82.2 | ±4.3 | 1 | terminalbench | SINGLE |
| 11 | GPT-5.5 | codex | 82.0 | ±4.3 | 1 | terminalbench | SINGLE |
| 12 | GPT-5.4 (2026-03-05) | forgecode | 81.8 | ±3.9 | 1 | terminalbench | SINGLE |
| 13 | Pointer Agent W/ Sonnet 4.6 | agentic-framework / 100-steps | 81.5 | — | 1 | osworld | SINGLE |
| 14 | Muse Spark 1.1 | general-model / 100-steps | 80.7 | — | 1 | osworld | SINGLE |
| 15 | Claude Opus 4.7 | wozcode | 80.2 | ±4.1 | 1 | terminalbench | SINGLE |
| # | Model | Head to head | Played | Benches | Best finish | Overall |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 97.8% | 91 of 93 | 3 | 1 of 23 on DeepSWE | 1 |
| 2 | GPT-5.6 Sol | 95.7% | 89 of 93 | 3 | 1 of 82 on ALE-Bench | 2 |
| 3 | Claude Fable 5 | 94.6% | 88 of 93 | 3 | 1 of 25 on FrontierCode | 3 |
| 4 | Kimi K3 | 82.8% | 77 of 93 | 3 | 8 of 82 on ALE-Bench | 4 |
| 5 | Grok 4.6 | 81.7% | 76 of 93 | 3 | 3 of 25 on FrontierCode | 5 |
| 6 | GPT-5.5 | 78.5% | 73 of 93 | 3 | 7 of 82 on ALE-Bench | 10 |
| 7 | GPT-5.6 Terra | 76.9% | 71.5 of 93 | 3 | 4 of 82 on ALE-Bench | 12 |
| 8 | GLM-5.3 | 75.7% | 53 of 70 | 2 | 3 of 23 on DeepSWE | 8 |
| 9 | Claude Opus 4.8 | 72.6% | 67.5 of 93 | 3 | 12 of 82 on ALE-Bench | 6 |
| 10 | GPT-5.6 Luna | 68.8% | 64 of 93 | 3 | 5 of 82 on ALE-Bench | 16 |
Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 54 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.
This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.
275 systems with at least one qualifying benchmark: 16 FULL, 1 THIN, 258 SINGLE. 212 contributing measurements carry no measurement date. Showing 15.
clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.
A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.
| Row | System | Configuration | Score | Interval | Benches | Which | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5 (2025-08-07) | diff / high | 88.0 | — | 1 | aider_polyglot | SINGLE |
| 2 | GPT-5 (2025-08-07) | diff / medium | 86.7 | — | 1 | aider_polyglot | SINGLE |
| 3 | OpenAI o3-pro (2025-06-10) | diff / high | 84.9 | — | 1 | aider_polyglot | SINGLE |
| 4 | Claude Opus 4.7 | max | 83.5 | ±3.3 | 1 | swe_bench_verified | SINGLE |
| 5 | Gemini 2.5 Pro Preview 0605 | diff-fenced / 32k | 83.1 | — | 1 | aider_polyglot | SINGLE |
| 6 | GPT-5 (2025-08-07) | diff / low | 81.3 | — | 1 | aider_polyglot | SINGLE |
| 7 | o3 (2025-04-16) | diff / high | 81.3 | — | 1 | aider_polyglot | SINGLE |
| 8 | GPT-5.5 Pre Release | xhigh | 80.6 | ±3.5 | 1 | swe_bench_verified | SINGLE |
| 9 | Grok 4.0709 | diff | 79.6 | — | 1 | aider_polyglot | SINGLE |
| 10 | Grok 4.0709 | diff / high | 79.6 | — | 1 | aider_polyglot | SINGLE |
| 11 | Gemini 2.5 Pro Preview 0605 | diff-fenced | 79.1 | — | 1 | aider_polyglot | SINGLE |
| 12 | DeepSeek V4 Pro | max | 77.6 | ±3.7 | 1 | swe_bench_verified | SINGLE |
| 13 | Gemini 2.5 Pro Preview 0506 | diff-fenced | 76.9 | — | 1 | aider_polyglot | SINGLE |
| 14 | o3 (2025-04-16) | diff / medium | 76.9 | — | 1 | aider_polyglot | SINGLE |
| 15 | o3 (2025-04-16) | diff | 76.9 | — | 1 | aider_polyglot | SINGLE |
| # | Model | Head to head | Played | Benches | Best finish | Overall |
|---|---|---|---|---|---|---|
| 1 | GPT-5.5 Pre Release | 99.1% | 382.5 of 386 | 3 | 1 of 67 on FrontierMath | null |
| 2 | GPT-5.5 Pro Pre Release | 98.6% | 380.5 of 386 | 3 | 2 of 161 on OTIS Mock AIME 2024-2025 | null |
| 3 | GPT-5.4 Pro (2026-03-05) | 97.7% | 317.5 of 325 | 4 | 2 of 163 on GPQA Diamond (Epoch's own run) | null |
| 4 | Claude Opus 5 | 96.5% | 371.5 of 385 | 3 | 1 of 66 on ARC-AGI-2 | 1 |
| 5 | Qwen3.8 Max | 96.6% | 311 of 322 | 2 | 4 of 161 on OTIS Mock AIME 2024-2025 | null |
| 6 | Gemini 3.7 Flash | 94.8% | 365 of 385 | 3 | 1 of 163 on GPQA Diamond (Epoch's own run) | null |
| 7 | Grok 4.6 | 93.6% | 360.5 of 385 | 3 | 7 of 163 on GPQA Diamond (Epoch's own run) | 5 |
| 8 | Gemini 3.1 Pro Preview | 93.5% | 453.5 of 485 | 5 | 1 of 37 on Humanity's Last Exam | 17 |
| 9 | Grok 4.5 | 92.9% | 357.5 of 385 | 3 | 6 of 163 on GPQA Diamond (Epoch's own run) | 11 |
| 10 | DeepSeek V4 Pro 0813 | 92.5% | 356 of 385 | 3 | 5 of 161 on OTIS Mock AIME 2024-2025 | null |
Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 168 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.
This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.
398 systems with at least one qualifying benchmark: 265 FULL, 1 THIN, 132 SINGLE. 274 contributing measurements carry no measurement date. Showing 15.
clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.
A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.
| Row | System | Configuration | Score | Interval | Benches | Which | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 Max | xhigh | 96.1 | ±1.7 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 2 | Claude Opus 5 | as published | 95.4 | ±2.8 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 3 | GPT-5.6 Sol | max | 95.3 | ±1.0 | 3 | arc_agi_2, gpqa_diamond, otis_mock_aime | FULL |
| 4 | Gemini 3.1 Pro Preview | high | 95.0 | ±3.4 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 5 | Claude Opus 5 | max | 94.0 | ±1.3 | 3 | arc_agi_2, gpqa_diamond, otis_mock_aime | FULL |
| 6 | DeepSeek V4 Pro | high | 93.2 | ±3.6 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 7 | Qwen3.7 Max | as published | 93.2 | ±3.6 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 8 | DeepSeek V4 Pro | max | 93.2 | ±2.6 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 9 | Claude Sonnet 5 | xhigh | 92.6 | ±2.8 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 10 | Gemini 3 Flash Preview | high | 92.5 | ±3.7 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 11 | GPT-5.6 Terra | max | 92.3 | ±1.0 | 3 | arc_agi_2, gpqa_diamond, otis_mock_aime | FULL |
| 12 | Gemini 3.7 Flash | high | 92.2 | ±1.5 | 3 | arc_agi_2, gpqa_diamond, otis_mock_aime | FULL |
| 13 | GLM-5.3-Flash | max | 92.0 | ±2.7 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 14 | Kimi K2.7 Code | as published | 91.7 | ±3.8 | 2 | gpqa_diamond, otis_mock_aime | FULL |
| 15 | Claude Fable 5 | max | 91.6 | ±1.6 | 3 | arc_agi_2, gpqa_diamond, otis_mock_aime | FULL |
| # | Model | Head to head | Played | Benches | Best finish | Overall |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 97.4% | 187 of 192 | 2 | 1 of 158 on LMArena Text (style-controlled) | 3 |
| 2 | Claude Opus 5 High | 95.3% | 183 of 192 | 2 | 4 of 100 on LMArena WebDev | null |
| 3 | Claude Opus 5 Max | 94.8% | 182 of 192 | 2 | 1 of 100 on LMArena WebDev | null |
| 4 | Kimi K3 Max | 94.8% | 182 of 192 | 2 | 2 of 100 on LMArena WebDev | null |
| 5 | Claude Opus 4.7 High | 91.7% | 176 of 192 | 2 | 3 of 158 on LMArena Text (style-controlled) | null |
| 6 | Claude Opus 4.6 High | 91.1% | 175 of 192 | 2 | 2 of 158 on LMArena Text (style-controlled) | null |
| 7 | Gemini 3.7 Flash High | 91.1% | 175 of 192 | 2 | 9 of 158 on LMArena Text (style-controlled) | null |
| 8 | Claude Opus 4.7 | 90.6% | 174 of 192 | 2 | 6 of 158 on LMArena Text (style-controlled) | 9 |
| 9 | GLM-5.3 Max | 89.6% | 172 of 192 | 2 | 8 of 100 on LMArena WebDev | null |
| 10 | Qwen3.8 Max | 89.6% | 172 of 192 | 2 | 3 of 100 on LMArena WebDev | null |
Ranked head to head: ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. Showing 10 of 97 models with enough evidence in this classification. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.
This table is evidence, not a ranking. It is sorted by score, and score is not comparable across rows: a row resting on one benchmark sits beside a row resting on three, and the benchmarks themselves differ in difficulty by up to 58 normalised points. Read it to see how much a configuration changes a result, and read the table above for where a model actually places.
419 systems with at least one qualifying benchmark: 0 FULL, 93 THIN, 326 SINGLE. 512 contributing measurements carry no measurement date. Showing 15.
clamped means the fixed anchor cut the value off at the top or the bottom of its scale, so that score is a floor or a ceiling rather than a point. It is the signature of an anchor that needs widening, and it is flagged rather than smoothed away. sources disagree means two benchmarks in the same correlation family, which are supposed to measure the same thing, returned materially different answers, so the interval was widened rather than allowed to shrink.
A model name can appear on more than one row. Those rows are not duplicates: they are different configurations of the same model, and the Configuration column names what differs. One model run under four agent harnesses is four measured systems with four different scores, and averaging them away would hide how much of a result comes from the harness rather than the model.
| Row | System | Configuration | Score | Interval | Benches | Which | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 Max | as published | 85.2 | ±1.2 | 2 | arena_text_style_control, arena_webdev sources disagree |
THIN |
| 2 | Kimi K3 Max | as published | 84.2 | ±1.4 | 2 | arena_text_style_control, arena_webdev | THIN |
| 3 | Muse Spark | as published | 84.1 | ±0.8 | 1 | arena_text_style_control | SINGLE |
| 4 | Claude Opus 5 High | as published | 83.7 | ±1.0 | 2 | arena_text_style_control, arena_webdev | THIN |
| 5 | GPT-5.5 High | as published | 83.2 | ±0.6 | 1 | arena_text_style_control | SINGLE |
| 6 | Qwen3.8 Max | as published | 83.2 | ±1.6 | 2 | arena_text_style_control, arena_webdev | THIN |
| 7 | GPT-5.6 Sol Xhigh | as published | 83.2 | ±0.7 | 1 | arena_text_style_control | SINGLE |
| 8 | Claude Fable 5 | as published | 82.5 | ±4.3 | 2 | arena_text_style_control, arena_webdev sources disagree |
THIN |
| 9 | GPT-5.5 | as published | 82.4 | ±0.6 | 1 | arena_text_style_control | SINGLE |
| 10 | GPT-5.4 High | as published | 82.4 | ±0.6 | 1 | arena_text_style_control | SINGLE |
| 11 | GPT-5.2 Chat Latest (2026-02-10) | as published | 82.3 | ±0.6 | 1 | arena_text_style_control | SINGLE |
| 12 | Grok 4.20 Beta1 | as published | 82.1 | ±0.7 | 1 | arena_text_style_control | SINGLE |
| 13 | Qwen3.7 Max Preview | as published | 82.0 | ±1.4 | 1 | arena_text_style_control | SINGLE |
| 14 | GPT-5.5 Instant | as published | 81.9 | ±0.7 | 1 | arena_text_style_control | SINGLE |
| 15 | Grok 4.20 Multi Agent Beta 0309 | as published | 81.5 | ±0.5 | 1 | arena_text_style_control | SINGLE |
Who measured it, who moved it, on what terms, on what scale, and how many of its rows survived to enter a score.
| ID | Benchmark | Category | Evaluator | Transport | Licence | Metric | Anchors | Observed | Interval | Attribution | Contamination | Ingested | Scored |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| terminalbench | Terminal-Bench 2 | Agents | Terminal-Bench | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.03 to 0.85 | stderr | row | low | 212 | 104 |
| apex_agents | APEX Agents | Agents | APEX | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.01 to 0.45 | stderr | file | low | 63 | 62 |
| metr_time_horizon | METR time horizons | Agents | METR | Epoch AI Benchmarking Hub | CC-BY-4.0 | time_horizon_min | 0.00 to 3.50 | 0.05 to 1044.78 | ci | row | low | 50 | 36 |
| cybench | Cybench | Agents | Cybench | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.05 to 0.93 | none | row | medium | 22 | 16 |
| arena_agent | LMArena Agent | Agents | LMArena | LMArena leaderboard dataset | CC-BY-4.0 | ips_score | -0.05 to 0.25 | -0.21 to 0.13 | ci | own | n/a | 54 | 54 |
| swe_bench_verified | SWE-bench Verified (Epoch's own run) | Coding | Epoch AI | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.31 to 0.83 | stderr | own | medium | 41 | 40 |
| aider_polyglot | Aider Polyglot | Coding | Aider | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.04 to 0.88 | none | row | medium | 72 | 72 |
| deepswe | DeepSWE | Coding | Datacurve | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.02 to 0.74 | ci_halfwidth | row | medium | 62 | 61 |
| frontiercode | FrontierCode | Coding | Cognition | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.08 to 0.53 | none | row | low | 28 | 28 |
| gpqa_diamond | GPQA Diamond (Epoch's own run) | Reasoning | Epoch AI | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.13 to 0.95 | stderr | own | high | 284 | 280 |
| frontiermath | FrontierMath | Reasoning | Epoch AI | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.00 to 0.52 | stderr | own | low | 101 | 101 |
| hle | Humanity's Last Exam | Reasoning | Humanity’s Last Exam (CAIS / Scale AI) | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.03 to 0.46 | stderr | file | low | 59 | 53 |
| otis_mock_aime | OTIS Mock AIME 2024-2025 | Reasoning | Epoch AI | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.00 to 1.00 | stderr | own | medium | 256 | 256 |
| arc_agi_2 | ARC-AGI-2 | Reasoning | ARC Prize | Epoch AI Benchmarking Hub | CC-BY-4.0 | pass_rate | 0.00 to 1.00 | 0.00 to 0.93 | none | file | low | 216 | 216 |
| arena_text_style_control | LMArena Text (style-controlled) | Preference | LMArena | LMArena leaderboard dataset | CC-BY-4.0 | bt_rating | 900.00 to 1600.00 | 952.54 to 1507.48 | ci | own | n/a | 395 | 395 |
| arena_webdev | LMArena WebDev | Preference | LMArena | LMArena leaderboard dataset | CC-BY-4.0 | bt_rating | 1000.00 to 1800.00 | 1080.11 to 1690.64 | ci | own | n/a | 117 | 117 |
| ale_bench | ALE-Bench | Coding | ALE-Bench (Sakana AI with AtCoder) | Epoch AI Benchmarking Hub | CC-BY-4.0 | perf_rating | 0.00 to 3500.00 | 137.78 to 2176.88 | none | file | low | 110 | 110 |
| terminalbench_4 | Terminal-Bench 4.0 | Agents | Terminal-Bench | Terminal-Bench 4.0 leaderboard submissions | Apache-2.0 | pass_rate | 0.00 to 1.00 | — | ci_halfwidth | own | low | 10 | 10 |
| tau2_airline | Tau2-Bench Airline | Agents | Tau2-Bench (Sierra) | Tau2-Bench leaderboard | MIT | pass_rate | 0.00 to 1.00 | — | none | own | low | 0 | 0 |
| tau2_retail | Tau2-Bench Retail | Agents | Tau2-Bench (Sierra) | Tau2-Bench leaderboard | MIT | pass_rate | 0.00 to 1.00 | — | none | own | low | 0 | 0 |
| tau2_telecom | Tau2-Bench Telecom | Agents | Tau2-Bench (Sierra) | Tau2-Bench leaderboard | MIT | pass_rate | 0.00 to 1.00 | — | none | own | low | 0 | 0 |
| tau2_banking | Tau2-Bench Banking | Agents | Tau2-Bench (Sierra) | Tau2-Bench leaderboard | MIT | pass_rate | 0.00 to 1.00 | — | none | own | low | 0 | 0 |
| osworld | OSWorld | Agents | OSWorld (XLANG Lab) | OSWorld leaderboard | none stated | pass_rate | 0.00 to 1.00 | — | none | own | low | 164 | 110 |
| browsecomp | BrowseComp | Agents | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 7 | 1 |
| mcp_atlas | MCP-Atlas | Agents | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 4 | 0 |
| tool_decathlon | Tool-Decathlon | Agents | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 2 | 0 |
| cybergym | CyberGym | Agents | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 0 | 0 |
| swe_bench_pro | SWE-bench Pro | Coding | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 2 | 0 |
| swe_bench_multilingual | SWE-bench Multilingual | Coding | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 1 | 1 |
| aime | AIME | Reasoning | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 3 | 0 |
| hmmt | HMMT | Reasoning | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 4 | 0 |
| imo_answerbench | IMOAnswerBench | Reasoning | various; recorded per row | Model benchmark dossiers | facts, not expression | pass_rate | 0.00 to 1.00 | — | none | row | low | 4 | 2 |
Table scrolls sideways
Ingested and Scored are deliberately different numbers, and so is the row count in the source signature printed in section 11. A source file has some number of rows; some of those rows have no usable value or no identifiable model and never become an observation; and some observations are then refused entry to a score by the rights gate or the provenance gate. Three numbers, narrowing at each step. Section 07 accounts for the difference. Contamination is our recorded judgement of how likely a benchmark's tasks are to have appeared in training data. It is a risk flag, not a measurement, and a high rating is a reason to weigh that benchmark's contribution more carefully rather than to discard it.
Evaluator is who ran the benchmark. Transport is who we obtained it from. They are recorded separately and both are credited, because crediting only the aggregator would misstate who did the work. Attribution says how the evaluator is known: own means the transport ran the evaluation itself, row means the file carries a per-row source column so each row can be screened for vendor self-reports, and file means the file has no attribution column at all, so the evaluator is a reviewed assertion about the whole file and individual rows cannot be screened. Rows from a file-level source are counted and reported in the next section.
Benchmarks that measure the same thing are averaged inside their family before the category mean, so four variants of one task distribution count as one piece of evidence rather than four.
An exclusion the reader cannot see is the same as no exclusion at all.
| State | Rows | Why |
|---|---|---|
| UNATTRIBUTED | 114 | The file carries a source column and this row left it blank, so nobody can be named as the evaluator. |
| CLAIMED | 104 | A vendor self-report. Stored and displayable, never scored. |
| rights | 0 | The licence on the source does not permit us to restate its numbers. We may read it and link to it, and we do, but it never enters a score. |
| Measure | Count | What it means |
|---|---|---|
| Unscreened rows | 439 | From files with no attribution column. The evaluator is named for the file as a whole, but no individual row can be checked against it. |
| Unmapped identifiers | 836 | Model identifiers seen in the data that no reviewed alias covers. They are recorded rather than guessed at, because a wrong merge is worse than a missing row. |
| Metadata conflicts | 740 | Model keys where providers disagree about price, context window or open weights. Where they disagree we publish nothing for that field rather than picking one. |
| Systems seen but not ranked | 1145 | Identified, scored where possible, and short of the entry requirement in at least one required category. |
Every limit below is measured, not estimated, and every one of them is a reason to read a specific number with more care.
UNCALIBRATED - anchors and thresholds are placeholders until the 90-day backfill study runs. Every published number must carry this label. Anchors and thresholds are placeholders. Ordering within the board is far more trustworthy than the absolute value of any index.
A category score is the mean across its benchmark families, and those families are not equally hard. In Coding alone this report scores SWE-bench Verified (Epoch's own run), Aider Polyglot, DeepSWE, FrontierCode, ALE-Bench, SWE-bench Pro, SWE-bench Multilingual, whose normalised distributions sit at very different heights. Until the calibration study equalises them, being measured on a harder benchmark costs a model points that a reader would naturally attribute to weaker ability.
RE-MEASURED on 2026-08-29 when the index moved from composite category scores to head-to-head positions, because the previous band of 2.17 was derived in composite points and could not be carried onto a different scale. The measurement is the same one, run again: for the 13 ranked models holding both a qualifying preference position and all three required categories, compare the measured preference position with the position implied by the weighted mean of the others. THE BIAS DISAPPEARED. It was +21.66 composite points, 8 of 8 in the same direction. It is now -1.64 position points with an sd of 17.79 and only 5 of 13 positive, which is indistinguishable from zero. That is exactly what the previous version of this note predicted would happen: it argued the offset was an artefact of incommensurable anchors rather than a property of models, since preference is a Bradley-Terry rating normalised against placeholder anchors of its own while the other categories are pass rates against 0 to 1. Positions are scale-free, so the artefact had nothing left to live in. The hypothesis was recorded before the test and the test agreed with it. The band therefore now describes NOISE, not bias. Imputing preference moves a row by 17.79 position points one standard deviation, which at the 10% preference weight is 1.78 index points, so an ordering finer than that between an imputed row and a measured row is not something this board can assert. On the live snapshot the band contains 18 of 561 pairs, 10 of them with exactly one side imputed, against a median adjacent gap of 1.95 index points.
There is no longer an offset to correct, and that is the finding rather than a convenience. The previous version of this field declined to add back a +21.66 point offset on the grounds that it was an artefact of incommensurable anchors rather than a property of models: preference is a Bradley-Terry rating normalised against placeholder anchors of its own, 900 to 1600 for the text arena and 1000 to 1800 for the web development arena, which run on different scales and so cannot share one, while the other categories are pass rates against 0 to 1. Moving the index to positions removed the anchors from the comparison, and the offset went with them. The band is now a width and not a shift, so there is nothing to add back and adding anything would be fitting noise.
Measured on 2026-08-29 over 13 models. That is a thin sample, and it is re-measured as coverage grows.
439 scored rows come from files that carry no attribution column. The evaluator is named for the file, and the file is reviewed, but no individual row can be checked. They are marked rather than dropped, because dropping them would remove real coverage and hide the fact that the source does not publish attribution.
1278 contributing measurements across the category boards were published without a measurement date. They are counted and shown rather than assumed current or silently dropped.
A model can be missing because no one has published two independent measurements of it in a required category. That is a statement about public benchmark coverage, not about the model.
Every figure here traces to a snapshot id and a source manifest id, both printed at the top of this report and repeated in the footer. Quote them when reporting a problem and the exact inputs can be replayed to the byte. The same snapshot is available as a machine-readable document, so any figure on this page can be checked against the object it was rendered from without taking our word for it.
Corrections go to ferroxlabs.com. A correction that changes a published number is described when it is applied, not quietly folded into the next snapshot.
Required on any page that shows these numbers.
| Source | Status | Observations | Signature |
|---|---|---|---|
| arena | ok | 566 | arena_agent:54,arena_text_style_control:395,arena_webdev:117 |
| dossier | ok | 61 | dossiers:3,kept:61,dropped:178,hash:c9db018492921a02 |
| epoch | ok | 1546 | aider_polyglot:77,ale_bench:110,apex_agents:62,arc_agi_2:216,cybench:22,deepswe:61,frontiercode:33,frontiermath:101,gpqa_diamond:280,hle:51,metr_time_horizon:50,otis_mock_aime:256,swe_bench_verified:35,terminalbench:204 |
| modelsdev | ok | 0 | providers:211,models:7488 |
| osworld | ok | 160 | verified:1076,self:56,kept:160,dropped:972,hash:65022c7f3262daac |
| tbench4 | ok | 10 | submissions:10,kept:10,dropped:0,tree:de3608c2dd501f81 |