THE FORGE RANKINGSupdated 2026-08-30

One benchmark is one exam.
This is the whole field.

37 models scored across 32 independently run benchmarks. Every number names who measured it and links to where you can read it. Nothing here was run by us, and no vendor self-report is scored.

The field 37 ranked
23.1Claude 4 Sonnet median 65.5 99.5Claude Opus 5
37 models ranked
32 benchmarks
2,343 measurements
31% of the grid measured
0 self-reports scored104 shown

Leading right now

see the full board

Category leaders rank head to head: every model is compared to every other only on the benchmarks they BOTH sat, and the figure is the share of those comparisons it won. Raw scores are not comparable across models, because the benchmarks are neither equally hard nor sat by equally strong fields. How this is worked out.

The Ferrox Index

Current rankings

99.523.1 index range, 37 models

Bar length is the index: a model's position in each category, weighted and averaged. Each band inside a bar is what that category contributed, so you can see whether a lead was won on reasoning or on agents. A category counts for where a model finished against the others who sat the same benchmarks, never for the raw score it posted on them.

Agents35%Coding30%Reasoning25%Preference10%
1
A+
Claude Opus 5 Anthropic
99.5
2
A+
95.5
3
A+
Claude Fable 5 Anthropic
94.8
4
A+
Kimi K3 Moonshot
93.7
5
A
92.8
6
A
Claude Opus 4.8 Anthropic
87.1
7
A
85.5
8
A
GLM-5.3 Z.ai (Zhipu AI)
84.9
9
A
Claude Opus 4.7 Anthropic
84.0
10
A
GPT-5.5 OpenAI
81.6
11
A
81.0
12
A
80.1
13
B
Claude Sonnet 5 Anthropic
74.2
14
B
GLM-5.2 Z.ai (Zhipu AI)
71.1
15
B
Claude Opus 4.6 Anthropic
70.8
16
B
70.8
17
B
Gemini 3.1 Pro Preview Google DeepMind
69.9
18
C
Kimi K2.6 Moonshot
65.9
19
C
65.5
20
C
Gemini 3 Pro Preview Google DeepMind
65.3
21
C
Gemini 3 Flash Preview Google DeepMind
65.0
22
C
62.6
23
C
61.5
24
C
Claude 4.5 Opus Anthropic
61.4
25
D
GLM 5.1 Z.ai (Zhipu AI)
58.7
26
D
47.4
27
D
45.2
28
D
Kimi K2.5 Moonshot
43.4
29
D
42.2
30
D
MiniMax-M3 MiniMax
40.6
31
D
38.8
32
E
Claude 4.1 Opus Anthropic
34.9
33
E
Claude 4 Opus Anthropic
30.0
34
E
Gemini 2.5 Pro Google DeepMind
27.7
35
E
24.1
36
E
23.3
37
E
Claude 4 Sonnet Anthropic
23.1

Showing 37 of 37 ranked models. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37. Every name links to that model's own page, with the field behind each measurement and a ledger of who ran it.

Tops by category

Where each model actually ranks

63point swing, best to worst category

Overall rank is one number and it hides the interesting part. Each cell is this model's rank within that classification, out of every ranked model with enough evidence in it. Ordering is head to head on shared benchmarks only: a model should not gain from sitting easier exams, nor from sitting them against a weaker field, and under the earlier rules it did both. The widest swing on this board is GPT-OSS 120B, 63 percentile points between its best category and its worst. Hover a cell for the record behind it.

#ModelGrade AgentsCodingReasoningPreferenceEvidence
1 Claude Opus 5 A+ 1/54 1/54 4/168no data 9/32
2 GPT-5.6 Sol A+ 4/54 2/54 11/168no data 9/32
3 Claude Fable 5 A+ 2/54 3/54 24/168 1/97 11/32
4 Kimi K3 A+ 3/54 4/54 19/168no data 8/32
5 Grok 4.6 A 6/54 5/54 7/168no data 9/32
6 Claude Opus 4.8 A 8/54 9/54 14/168 18/97 12/32
7 GPT-5.4 (2026-03-05) A 9/54 11/54 15/168no data 11/32
8 GLM-5.3 A 11/54 8/54 21/168no data 6/32
9 Claude Opus 4.7 A 7/54 13/54 31/168 8/97 13/32
10 GPT-5.5 A 5/54 6/54 65/168 33/97 11/32
11 Grok 4.5 A 13/54 15/54 9/168 20/97 11/32
12 GPT-5.6 Terra A 16/54 7/54 32/168no data 8/32
13 Claude Sonnet 5 B 12/54 12/54 66/168no data 8/32
14 GLM-5.2 B 10/54 25/54 44/168no data 9/32
15 Claude Opus 4.6 B 15/54 27/54 28/168 12/97 13/32
16 GPT-5.6 Luna B 19/54 10/54 63/168no data 8/32
17 Gemini 3.1 Pro Preview B 18/54 28/54 8/168 25/97 13/32
18 Kimi K2.6 C 29/54 19/54 13/168 36/97 10/32
19 GPT-5.2 (2025-12-11) C 22/54 24/54 29/168no data 9/32
20 Gemini 3 Pro Preview C 20/54 29/54 20/168no data 9/32
21 Gemini 3 Flash Preview C 27/54 16/54 40/168no data 8/32
22 Claude Sonnet 4.6 C 25/54 22/54 50/168 23/97 14/32
23 Kimi K2.7 Code C 23/54 31/54 22/168no data 8/32
24 Claude 4.5 Opus C 24/54 23/54 51/168 34/97 12/32
25 GLM 5.1 D 26/54 33/54 25/168 31/97 13/32
26 GPT-5.1 (2025-11-13) D 32/54 30/54 71/168no data 9/32
27 GPT-5 (2025-08-07) D 34/54 38/54 45/168no data 11/32
28 Kimi K2.5 D 39/54 34/54 49/168no data 6/32
29 o3 (2025-04-16) D 36/54 37/54 58/168no data 11/32
30 MiniMax-M3 D 30/54 41/54 88/168 45/97 9/32
31 Claude Sonnet 4.5 D 33/54 43/54 73/168 54/97 13/32
32 Claude 4.1 Opus E 37/54 42/54 82/168 59/97 11/32
33 Claude 4 Opus E 38/54 45/54 92/168no data 10/32
34 Gemini 2.5 Pro E 47/54 48/54 53/168 73/97 10/32
35 GPT-OSS 120B E 48/54 52/54 57/168no data 8/32
36 Claude 3.7 Sonnet E 44/54 49/54 91/168no data 11/32
37 Claude 4 Sonnet E 43/54 51/54 89/168no data 11/32

Evidence is benchmarks measured out of 32. Denominators differ by classification because they count only the models with enough evidence in that one. a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels.

What changed

rebuilt 2026-08-30

The rules changed between these two snapshots (benchmarks added: aime, browsecomp, cybergym, hmmt, imo_answerbench, mcp_atlas, swe_bench_multilingual, swe_bench_pro, tool_decathlon). New evidence is listed below, because that is a fact. Rank movement is not, because a model that moved when the scorer or the instrument moved did not move. The trend resumes from the next snapshot built under the current rules.

New benchmarks

  • AIME 0 of 37 measured
  • BrowseComp 1 of 37 measured
  • CyberGym 0 of 37 measured
  • HMMT 0 of 37 measured
  • IMOAnswerBench 1 of 37 measured
  • MCP-Atlas 0 of 37 measured
  • SWE-bench Multilingual 1 of 37 measured
  • SWE-bench Pro 0 of 37 measured
  • Tool-Decathlon 0 of 37 measured

Newly ranked

New measurements

The fields

Every model with evidence in a category, ranked in it

373ranked places across 4 fields

A coding board needs coding evidence, not all three categories. These boards hold every model with 2 qualifying benchmarks in that one classification, whether or not it has enough spread to carry an overall score. Ordering is head to head on shared benchmarks only. Most of these models are not missing from our reading of the record, they are missing from the world's agent and coding evaluations, which is why so many rank here and nowhere else.

Agents

54 models with qualifying evidence · 37 of them hold an overall score

#ModelOrg Win shareHead to headBenchmarks Overall
1 Claude Opus 5 Anthropic 98.6% 70/71 3 1
2 Claude Fable 5 Anthropic 97.2% 69/71 3 3
3 Kimi K3 Moonshot 87.1% 54/62 2 4
4 GPT-5.6 Sol OpenAI 83.1% 59/71 3 2
5 GPT-5.5 OpenAI 86.4% 76/88 3 10
6 Grok 4.6 xAI 74.6% 53/71 3 5
7 Claude Opus 4.7 Anthropic 79.5% 70/88 3 9
8 Claude Opus 4.8 Anthropic 71.8% 51/71 3 6
9 GPT-5.4 (2026-03-05) OpenAI 86.7% 72/83 3 7
10 GLM-5.2 Z.ai (Zhipu AI) 72.6% 45/62 2 14
11 GLM-5.3 Z.ai (Zhipu AI) 59.4% 19/32 2 8
12 Claude Sonnet 5 Anthropic 64.1% 45.5/71 3 13
13 Grok 4.5 xAI 64.1% 45.5/71 3 11
14 Muse Spark 1.1 no overall score Meta AI 69.1% 47/68 3
15 Claude Opus 4.6 Anthropic 67.0% 59/88 3 15
16 GPT-5.6 Terra OpenAI 43.8% 14/32 2 12
17 GPT-5.3 Codex no overall score OpenAI 70.8% 46/65 2
18 Gemini 3.1 Pro Preview Google DeepMind 67.0% 71/106 4 17
19 GPT-5.6 Luna OpenAI 40.6% 13/32 2 16
20 Gemini 3 Pro Preview Google DeepMind 64.6% 42/65 2 20
21 GPT 5.1 Codex Max no overall score OpenAI 75.0% 33/44 2
22 GPT-5.2 (2025-12-11) OpenAI 60.0% 39/65 2 19
23 Kimi K2.7 Code Moonshot 46.8% 29/62 2 23
24 Claude 4.5 Opus Anthropic 59.6% 49.5/83 3 24
25 Claude Sonnet 4.6 Anthropic 47.9% 45/94 4 22
26 GLM 5.1 Z.ai (Zhipu AI) 42.9% 21/49 2 25
27 Gemini 3 Flash Preview Google DeepMind 52.3% 34/65 2 21
28 GPT-5.1 Codex no overall score OpenAI 50.0% 32.5/65 2
29 Kimi K2.6 Moonshot 36.8% 25/68 3 18
30 MiniMax-M3 MiniMax 24.1% 7/29 2 30
31 GPT-5 Codex no overall score OpenAI 41.5% 27/65 2
32 GPT-5.1 (2025-11-13) OpenAI 38.5% 25/65 2 26
33 Claude Sonnet 4.5 Anthropic 51.7% 30/58 4 31
34 GPT-5 (2025-08-07) OpenAI 41.0% 34/83 3 27
35 MiniMax-M2.7 no overall score MiniMax 21.4% 10.5/49 2
36 o3 (2025-04-16) OpenAI 38.6% 22/57 2 29
37 Claude 4.1 Opus Anthropic 46.2% 24/52 3 32
38 Claude 4 Opus Anthropic 57.7% 15/26 2 33
39 Kimi K2.5 Moonshot 26.7% 12/45 2 28
40 Grok 4.0709 no overall score xAI 30.1% 25/83 3
41 Nemotron 3 Ultra (free) no overall score NVIDIA 15.3% 9.5/62 2
42 MiniMax-M2.5 no overall score MiniMax 21.5% 14/65 2
43 Claude 4 Sonnet Anthropic 32.3% 21/65 3 37
44 Claude 3.7 Sonnet Anthropic 34.4% 11/32 3 36
45 Claude Haiku 4.5 no overall score Anthropic 16.9% 11/65 2
46 Kimi K2 Thinking no overall score Moonshot 19.3% 16/83 3
47 Gemini 2.5 Pro Google DeepMind 10.8% 7/65 2 34
48 GPT-OSS 120B OpenAI 10.8% 9/83 3 35
49 Claude 3.5 Sonnet (2024-06-20) no overall score Anthropic 19.2% 5/26 2
50 Z-AI/GLM 4.6 no overall score Z.ai (Zhipu AI),Tsinghua University 6.2% 4/65 2
51 o1 (2024-12-17) no overall score OpenAI 7.9% 4.5/57 2
52 o1 Preview (2024-09-12) no overall score OpenAI 13.5% 3.5/26 2
53 GPT-4o (2024-11-20) no overall score OpenAI 5.4% 3.5/65 3
54 Claude 3 Opus (2024-02-29) no overall score Anthropic 1.9% 0.5/26 2

Coding

54 models with qualifying evidence · 37 of them hold an overall score

#ModelOrg Win shareHead to headBenchmarks Overall
1 Claude Opus 5 Anthropic 97.8% 91/93 3 1
2 GPT-5.6 Sol OpenAI 95.7% 89/93 3 2
3 Claude Fable 5 Anthropic 94.6% 88/93 3 3
4 Kimi K3 Moonshot 82.8% 77/93 3 4
5 Grok 4.6 xAI 81.7% 76/93 3 5
6 GPT-5.5 OpenAI 78.5% 73/93 3 10
7 GPT-5.6 Terra OpenAI 76.9% 71.5/93 3 12
8 GLM-5.3 Z.ai (Zhipu AI) 75.7% 53/70 2 8
9 Claude Opus 4.8 Anthropic 72.6% 67.5/93 3 6
10 GPT-5.6 Luna OpenAI 68.8% 64/93 3 16
11 GPT-5.4 (2026-03-05) OpenAI 76.0% 74.5/98 3 7
12 Claude Sonnet 5 Anthropic 67.7% 63/93 3 13
13 Claude Opus 4.7 Anthropic 74.5% 76/102 3 9
14 GPT-5.3 Codex no overall score OpenAI 78.5% 62/79 2
15 Grok 4.5 xAI 64.0% 59.5/93 3 11
16 Gemini 3 Flash Preview Google DeepMind 73.4% 58/79 2 21
17 Gemini 3.7 Flash no overall score Google DeepMind 54.8% 51/93 3
18 Qwen3.7 Max no overall score Alibaba 69.6% 55/79 2
19 Kimi K2.6 Moonshot 61.4% 48.5/79 2 18
20 DeepSeek V4 Flash 0731 no overall score DeepSeek 51.4% 38/74 2
21 GPT-5.4 mini (2026-03-17) no overall score OpenAI 51.4% 38/74 2
22 Claude Sonnet 4.6 Anthropic 51.2% 62/121 4 22
23 Claude 4.5 Opus Anthropic 60.1% 47.5/79 2 24
24 GPT-5.2 (2025-12-11) OpenAI 60.1% 47.5/79 2 19
25 GLM-5.2 Z.ai (Zhipu AI) 50.4% 61/121 4 14
26 Gemini 3.5 Flash no overall score Google 52.0% 51/98 3
27 Claude Opus 4.6 Anthropic 50.0% 51/102 3 15
28 Gemini 3.1 Pro Preview Google DeepMind 40.0% 28/70 2 17
29 Gemini 3 Pro Preview Google DeepMind 50.6% 40/79 2 20
30 GPT-5.1 (2025-11-13) OpenAI 48.1% 38/79 2 26
31 Kimi K2.7 Code Moonshot 31.2% 29/93 3 23
32 Inkling no overall score Thinking Machines 32.4% 24/74 2
33 GLM 5.1 Z.ai (Zhipu AI) 39.2% 31/79 2 25
34 Kimi K2.5 Moonshot 38.6% 30.5/79 2 28
35 Gemini 3.6 Flash no overall score Google DeepMind 25.8% 24/93 3
36 o4 mini (2025-04-16) no overall score OpenAI 40.0% 24/60 2
37 o3 (2025-04-16) OpenAI 38.6% 34/88 3 29
38 GPT-5 (2025-08-07) OpenAI 37.5% 33/88 3 27
39 DeepSeek V4 Pro no overall score DeepSeek 28.4% 29/102 3
40 DeepSeek-R1-0528 no overall score DeepSeek 33.3% 20/60 2
41 MiniMax-M3 MiniMax 22.5% 23/102 3 30
42 Claude 4.1 Opus Anthropic 25.3% 20/79 2 32
43 Claude Sonnet 4.5 Anthropic 25.3% 20/79 2 31
44 Z-Ai/GLM 5 no overall score Z.ai (Zhipu AI) 25.3% 20/79 2
45 Claude 4 Opus Anthropic 32.4% 12/37 2 33
46 GPT-5 mini (2025-08-07) no overall score OpenAI 22.8% 18/79 2
47 Mistral Medium 3.5 no overall score Mistral AI 12.2% 9/74 2
48 Gemini 2.5 Pro Google DeepMind 15.2% 12/79 2 34
49 Claude 3.7 Sonnet Anthropic 18.9% 7/37 2 36
50 Qwen3.6 Plus no overall score Alibaba 10.1% 8/79 2
51 Claude 4 Sonnet Anthropic 11.7% 7/60 2 37
52 GPT-OSS 120B OpenAI 5.0% 3/60 2 35
53 Kimi K2 0905 no overall score Moonshot 5.0% 3/60 2
54 GPT-4.1 (2025-04-14) no overall score OpenAI 3.4% 3/88 3

Reasoning

168 models with qualifying evidence · 37 of them hold an overall score

#ModelOrg Win shareHead to headBenchmarks Overall
1 GPT-5.5 Pre Release no overall score OpenAI 99.1% 382.5/386 3
2 GPT-5.5 Pro Pre Release no overall score OpenAI 98.6% 380.5/386 3
3 GPT-5.4 Pro (2026-03-05) no overall score OpenAI 97.7% 317.5/325 4
4 Claude Opus 5 Anthropic 96.5% 371.5/385 3 1
5 Qwen3.8 Max no overall score Alibaba 96.6% 311/322 2
6 Gemini 3.7 Flash no overall score Google DeepMind 94.8% 365/385 3
7 Grok 4.6 xAI 93.6% 360.5/385 3 5
8 Gemini 3.1 Pro Preview Google DeepMind 93.5% 453.5/485 5 17
9 Grok 4.5 xAI 92.9% 357.5/385 3 11
10 DeepSeek V4 Pro 0813 no overall score DeepSeek 92.5% 356/385 3
11 GPT-5.6 Sol OpenAI 91.6% 352.5/385 3 2
12 Qwen3.7 Max no overall score Alibaba 91.8% 295.5/322 2
13 Kimi K2.6 Moonshot 91.3% 352.5/386 3 18
14 Claude Opus 4.8 Anthropic 90.2% 405/449 4 6
15 GPT-5.4 (2026-03-05) OpenAI 89.0% 431.5/485 5 7
16 DeepSeek V4 Pro no overall score DeepSeek 89.8% 289/322 2
17 GLM-5.3-Flash no overall score Z.ai (Zhipu AI) 89.3% 287.5/322 2
18 DeepSeek V4 Flash 0731 no overall score DeepSeek 88.1% 339/385 3
19 Kimi K3 Moonshot 87.9% 338.5/385 3 4
20 Gemini 3 Pro Preview Google DeepMind 85.6% 415/485 5 20
21 GLM-5.3 Z.ai (Zhipu AI) 86.6% 279/322 2 8
22 Kimi K2.7 Code Moonshot 86.3% 278/322 2 23
23 Muse Spark no overall score Meta AI 85.8% 362/422 4
24 Claude Fable 5 Anthropic 84.8% 326.5/385 3 3
25 GLM 5.1 Z.ai (Zhipu AI) 85.5% 361/422 4 25
26 Grok 4.3 no overall score xAI 85.9% 276.5/322 2
27 Grok 4.20 (Reasoning) no overall score xAI 85.4% 275/322 2
28 Claude Opus 4.6 Anthropic 83.7% 406/485 5 15
29 GPT-5.2 (2025-12-11) OpenAI 82.3% 399/485 5 19
30 Qwen3.6 Plus no overall score Alibaba 82.4% 318/386 3
31 Claude Opus 4.7 Anthropic 80.0% 388/485 5 9
32 GPT-5.6 Terra OpenAI 79.5% 306/385 3 12
33 Gemini 3.5 Flash no overall score Google 79.0% 354.5/449 4
34 Kimi K2p5 no overall score Moonshot 79.8% 308/386 3
35 Inkling no overall score Thinking Machines 78.6% 302.5/385 3
36 Inkling Small no overall score Thinking Machines 77.0% 296.5/385 3
37 Qwen3.6 Max Preview no overall score Alibaba 77.7% 300/386 3
38 GPT-5 Pro (2025-10-06) no overall score OpenAI 69.7% 69/99 2
39 Nemotron 3 Ultra (free) no overall score NVIDIA 73.9% 238/322 2
40 Gemini 3 Flash Preview Google DeepMind 70.8% 318/449 4 21
41 Gemini 3.6 Flash no overall score Google DeepMind 71.0% 273.5/385 3
42 GPT-5.4 mini (2026-03-17) no overall score OpenAI 70.3% 315.5/449 4
43 Grok 4.0709 no overall score xAI 70.2% 315/449 4
44 GLM-5.2 Z.ai (Zhipu AI) 70.3% 270.5/385 3 14
45 GPT-5 (2025-08-07) OpenAI 69.4% 336.5/485 5 27
46 Qwen3.5 Plus no overall score Alibaba 71.4% 275.5/386 3
47 DeepSeek Reasoner no overall score DeepSeek 72.0% 232/322 2
48 Qwen3.5 397B A17B no overall score Alibaba 71.4% 230/322 2
49 Kimi K2.5 Moonshot 61.6% 61/99 2 28
50 Claude Sonnet 4.6 Anthropic 67.9% 305/449 4 22
51 Claude 4.5 Opus Anthropic 66.6% 323/485 5 24
52 Z-Ai/GLM 5 no overall score Z.ai (Zhipu AI) 66.7% 299.5/449 4
53 Gemini 2.5 Pro Google DeepMind 65.9% 296/449 4 34
54 Kimi K2 Thinking Turbo no overall score Moonshot 69.3% 223/322 2
55 Gemini 2.5 Pro Experimental 0325 no overall score Google DeepMind 66.9% 132.5/198 2
56 Qwen3.6 Flash no overall score Alibaba 65.7% 253.5/386 3
57 GPT-OSS 120B OpenAI 65.4% 210.5/322 2 35
58 o3 (2025-04-16) OpenAI 60.9% 295.5/485 5 29
59 Qwen3.5 Flash no overall score Alibaba 63.0% 243/386 3
60 Qwen3 235B A22B Thinking 2507 no overall score Alibaba 62.8% 242.5/386 3
61 Qwen3.7 Plus no overall score Alibaba 63.4% 204/322 2
62 Qwen3.6 27B no overall score Alibaba 62.9% 202.5/322 2
63 GPT-5.6 Luna OpenAI 59.7% 230/385 3 16
64 Qwen3.6 35B A3B no overall score Alibaba 62.7% 202/322 2
65 GPT-5.5 OpenAI 59.2% 228/385 3 10
66 Claude Sonnet 5 Anthropic 62.1% 200/322 2 13
67 Z-Ai/GLM 4.7 no overall score Z.ai (Zhipu AI) 60.5% 233.5/386 3
68 Qwen3.7 Flash no overall score Alibaba 61.6% 198.5/322 2
69 Gemini 2.5 Pro Preview 0605 no overall score Google DeepMind 54.0% 175.5/325 4
70 o4 mini (2025-04-16) no overall score OpenAI 56.0% 271.5/485 5
71 GPT-5.1 (2025-11-13) OpenAI 55.7% 270/485 5 26
72 GPT-5.5 Instant no overall score OpenAI 59.6% 192/322 2
73 Claude Sonnet 4.5 Anthropic 54.2% 263/485 5 31
74 GPT-5 mini (2025-08-07) no overall score OpenAI 54.1% 262.5/485 5
75 Gemini 2.5 Flash Preview no overall score Google DeepMind 56.6% 111/196 2
76 Gemma 4 26B A4B IT no overall score Google DeepMind 58.1% 187/322 2
77 GPT-5.4 Nano (2026-03-17) no overall score OpenAI 53.6% 240.5/449 4
78 Gemini 2.5 Flash 0520 no overall score Google DeepMind 51.4% 133/259 3
79 Gemma 4 31B IT no overall score Google DeepMind 56.4% 181.5/322 2
80 o1 (2024-12-17) no overall score OpenAI 53.2% 224.5/422 4
81 Qwen3 Max (2025-09-23) no overall score Alibaba 54.2% 174.5/322 2
82 Claude 4.1 Opus Anthropic 50.1% 211.5/422 4 32
83 o3 mini (2025-01-31) no overall score OpenAI 48.2% 216.5/449 4
84 DeepSeek-R1-0528 no overall score DeepSeek 48.8% 188/385 3
85 Gemini 3.5 Flash Lite no overall score Google DeepMind 47.9% 184.5/385 3
86 Qwen3.5 9B no overall score Alibaba 49.4% 159/322 2
87 Grok 3 mini Beta no overall score xAI 46.9% 181/386 3
88 MiniMax-M3 MiniMax 47.8% 154/322 2 30
89 Claude 4 Sonnet Anthropic 41.6% 202/485 5 37
90 Gemini 2.5 Pro Preview 0506 no overall score Google DeepMind 43.9% 87/198 2
91 Claude 3.7 Sonnet Anthropic 41.2% 200/485 5 36
92 Claude 4 Opus Anthropic 41.0% 199/485 5 33
93 Gemini 3.1 Flash Lite no overall score Google 45.0% 161/358 3
94 DeepSeek-R1 no overall score DeepSeek 43.1% 166/385 3
95 DeepSeek Chat no overall score DeepSeek 45.3% 146/322 2
96 Grok 3 Beta no overall score xAI 42.7% 165/386 3
97 Qwq 32B no overall score Alibaba 44.7% 144/322 2
98 Gemini 2.0 Flash Thinking Exp 01.21 no overall score Google DeepMind,Google 40.1% 143.5/358 3
99 Claude Haiku 4.5 no overall score Anthropic 36.9% 165.5/449 4
100 DeepSeek-V3-0324 no overall score DeepSeek 41.5% 133.5/322 2
101 GPT-4.1 (2025-04-14) no overall score OpenAI 35.1% 170/485 5
102 GPT-4.5 Preview (2025-02-27) no overall score OpenAI 35.9% 151/421 4
103 OpenAI GPT-4.1 Mini no overall score OpenAI 35.7% 160.5/449 4
104 GPT-5 Nano (2025-08-07) no overall score OpenAI 35.7% 160.5/449 4
105 Gemini 2.0 Flash 001 no overall score Google DeepMind,Google 34.2% 153.5/449 4
106 o1 mini (2024-09-12) no overall score OpenAI 34.2% 153.5/449 4
107 R1 Distill Llama 70B no overall score DeepSeek 39.1% 126/322 2
108 GPT-OSS 20B no overall score OpenAI 37.6% 121/322 2
109 Gemini 1.5 Pro 002 no overall score Google DeepMind 30.3% 127.5/421 4
110 Llama 4 Maverick 17b 128e Instruct no overall score Meta AI 11.1% 11/99 2
111 Llama 4 Maverick 17B FP8 no overall score Meta AI 33.0% 127.5/386 3
112 Mistral Medium 3 no overall score Mistral AI 30.8% 130/422 4
113 Magistral Small 2506 no overall score Mistral AI 30.0% 115.5/385 3
114 DeepSeek-V3 no overall score DeepSeek 30.2% 116.5/386 3
115 o1 Preview (2024-09-12) no overall score OpenAI 32.1% 103.5/322 2
116 Qwen2.5-Max-2025-01-25 no overall score Alibaba 29.3% 113/386 3
117 Phi 4 no overall score Microsoft Research 30.9% 99.5/322 2
118 Claude 3.5 Sonnet (2024-10-22) no overall score Anthropic 25.9% 109.5/422 4
119 GPT-4.1 Nano (2025-04-14) no overall score OpenAI 24.7% 111/449 4
120 Grok 2.1212 no overall score xAI 26.3% 101.5/386 3
121 Gemma 3.27B It no overall score Google DeepMind 28.0% 90/322 2
122 Llama 3.1 405B Instruct no overall score Meta AI 27.6% 89/322 2
123 Qwen Plus (2025-01-25) no overall score Alibaba 27.6% 89/322 2
124 Qwen 3 8B no overall score Alibaba 27.2% 87.5/322 2
125 Claude 3.5 Sonnet (2024-06-20) no overall score Anthropic 24.1% 93/386 3
126 Mistral Large 2407 no overall score Mistral AI 25.9% 83.5/322 2
127 Qwen2.5 72B Instruct no overall score Alibaba 25.8% 83/322 2
128 Mistral Large 2411 no overall score Mistral AI 23.1% 89/386 3
129 Llama 4 Scout 17B 16E Instruct no overall score Meta AI 20.3% 91/449 4
130 Gemini 1.5 Flash 002 no overall score Google DeepMind 21.1% 81.5/386 3
131 GPT-4o (2024-08-06) no overall score OpenAI 20.3% 78.5/386 3
132 GPT-4o (2024-05-13) no overall score OpenAI 21.6% 69.5/322 2
133 Gemma 3 12B IT no overall score Google DeepMind 21.4% 69/322 2
134 GPT-4o (2024-11-20) no overall score OpenAI 15.5% 75/485 5
135 Qwen2.5 32B Instruct no overall score Alibaba 20.2% 65/322 2
136 GPT-4 Turbo (2024-04-09) no overall score OpenAI 19.9% 64/322 2
137 Mistral Small 2503 no overall score Mistral AI 19.4% 62.5/322 2
138 Gemini 1.5 Pro 001 no overall score Google DeepMind 19.3% 62/322 2
139 Llama 3.3 70B Instruct no overall score Meta AI 18.3% 59/322 2
140 Claude 3 Opus (2024-02-29) no overall score Anthropic 17.4% 56/322 2
141 GPT-4o-mini (2024-07-18) no overall score OpenAI 14.5% 56/385 3
142 Mistral Small 2501 no overall score Mistral AI 16.5% 53/322 2
143 Qwen Turbo (2024-11-01) no overall score Alibaba 16.5% 53/322 2
144 Llama 3.1 Tulu 3.70B Dpo no overall score Allen Institute for AI,University of Washington 16.1% 52/322 2
145 Llama 3.1 70b Instruct no overall score Meta AI 13.7% 44/322 2
146 Meta Llama 3.70B Instruct no overall score Meta AI 13.2% 42.5/322 2
147 Claude 3.5 Haiku (2024-10-22) no overall score Anthropic 10.9% 42/386 3
148 Llama-3.2-90B-Vision-Instruct no overall score Meta AI 12.7% 41/322 2
149 Gemini 1.5 Flash 001 no overall score Google DeepMind 12.4% 40/322 2
150 Claude 3 Sonnet (2024-02-29) no overall score Anthropic 11.8% 38/322 2
151 Gemma 3 4B IT no overall score Google DeepMind 11.8% 38/322 2
152 Gemini 1.5 Flash 8B 001 no overall score Google DeepMind 10.1% 32.5/322 2
153 Hermes 2 Theta Llama 3.70B no overall score Nous Research,Arcee AI 9.6% 31/322 2
154 Mistral Large 2402 no overall score Mistral AI 9.5% 30.5/322 2
155 Claude 2.0 no overall score Anthropic 8.4% 27/322 2
156 Anthropic: Claude 3 Haiku no overall score Anthropic 7.5% 24/322 2
157 Gemma 2 27B no overall score Google DeepMind 7.1% 23/322 2
158 Claude 2.1 no overall score Anthropic 6.5% 21/322 2
159 GPT-3.5 Turbo 0125 no overall score OpenAI 5.9% 19/322 2
160 Gemini 1.0 Pro 001 no overall score Google DeepMind 5.3% 17/322 2
161 GPT-4.0314 no overall score OpenAI 4.5% 14.5/322 2
162 GPT-4.0613 no overall score OpenAI 4.3% 14/322 2
163 Llama 3.1 8B Instruct no overall score Meta AI 4.3% 14/322 2
164 Gemma 2.9B It no overall score Google DeepMind 2.6% 8.5/322 2
165 Meta Llama 3.8B Instruct no overall score Meta AI 2.0% 6.5/322 2
166 Gemma 3.1B It no overall score Google DeepMind 1.9% 6/322 2
167 Deepseek Llm 67B Chat no overall score DeepSeek 1.7% 5.5/322 2
168 Llama 2.70B Chat Hf no overall score Meta AI 1.2% 4/322 2

Preference

97 models with qualifying evidence · 15 of them hold an overall score

#ModelOrg Win shareHead to headBenchmarks Overall
1 Claude Fable 5 Anthropic 97.4% 187/192 2 3
2 Claude Opus 5 High no overall score anthropic 95.3% 183/192 2
3 Claude Opus 5 Max no overall score anthropic 94.8% 182/192 2
4 Kimi K3 Max no overall score moonshot 94.8% 182/192 2
5 Claude Opus 4.7 High no overall score anthropic 91.7% 176/192 2
6 Claude Opus 4.6 High no overall score anthropic 91.1% 175/192 2
7 Gemini 3.7 Flash High no overall score google 91.1% 175/192 2
8 Claude Opus 4.7 Anthropic 90.6% 174/192 2 9
9 GLM-5.3 Max no overall score zai 89.6% 172/192 2
10 Qwen3.8 Max no overall score Alibaba 89.6% 172/192 2
11 GPT-5.6 Sol Xhigh no overall score openai 89.1% 171/192 2
12 Claude Opus 4.6 Anthropic 87.0% 167/192 2 15
13 Muse Spark 1.1 no overall score Meta AI 87.0% 167/192 2
14 Muse Spark 1.2 no overall score Meta AI 87.0% 167/192 2
15 Claude Opus 4.8 High no overall score anthropic 85.4% 164/192 2
16 Gemini 3.6 Flash High no overall score google 81.3% 156/192 2
17 GLM-5.2 Max no overall score zai 81.3% 156/192 2
18 Claude Opus 4.8 Anthropic 78.9% 151.5/192 2 6
19 Grok 4.6 High no overall score xai 78.6% 151/192 2
20 Grok 4.5 xAI 77.6% 149/192 2 11
21 Deepseek V4 Pro High (2026-08-13) no overall score deepseek 76.0% 146/192 2
22 GPT-5.5 High no overall score openai 75.5% 145/192 2
23 Claude Sonnet 4.6 Anthropic 75.0% 144/192 2 22
24 Gemini 3.5 Flash High no overall score google 75.0% 144/192 2
25 Gemini 3.1 Pro Preview Google DeepMind 72.4% 139/192 2 17
26 Gemini 3.5 Flash Medium no overall score google 72.4% 139/192 2
27 Claude Opus 4.5 20251101 High 32k no overall score anthropic 72.1% 138.5/192 2
28 Claude Sonnet 5 High no overall score anthropic 70.8% 136/192 2
29 Gemini 3 Pro Preview no overall score google 70.8% 136/192 2
30 GPT-5.6 Terra Xhigh no overall score openai 70.8% 136/192 2
31 GLM 5.1 Z.ai (Zhipu AI) 69.8% 134/192 2 25
32 GPT-5.4 High no overall score openai 69.3% 133/192 2
33 GPT-5.5 OpenAI 68.8% 132/192 2 10
34 Claude 4.5 Opus Anthropic 66.1% 127/192 2 24
35 MiMo V2.5 Pro (Crof) no overall score Xiaomi Corp 66.1% 127/192 2
36 Kimi K2.6 Moonshot 65.6% 126/192 2 18
37 Hy3 (free) no overall score tencent 63.5% 122/192 2
38 Qwen3.8 27B no overall score alibaba 62.2% 119.5/192 2
39 Qwen3.6 Max Preview no overall score Alibaba 62.0% 119/192 2
40 GPT-5.6 Luna Xhigh no overall score openai 61.5% 118/192 2
41 Deepseek V4 Pro High Preview no overall score deepseek 56.8% 109/192 2
42 Gemini 3.5 Flash Lite no overall score Google DeepMind 56.3% 108/192 2
43 DeepSeek V4 Pro no overall score DeepSeek 55.7% 107/192 2
44 Z-Ai/GLM 5 no overall score Z.ai (Zhipu AI) 54.2% 104/192 2
45 MiniMax-M3 MiniMax 53.6% 103/192 2 30
46 GPT-5.4 no overall score openai 51.6% 99/192 2
47 Grok 4.20 Beta Reasoning (0309) (xAI) no overall score xai 51.6% 99/192 2
48 Qwen3.6 Plus no overall score Alibaba 51.0% 98/192 2
49 Kimi K2.5 Thinking no overall score moonshot 50.0% 96/192 2
50 MiMo-V2-Pro no overall score Xiaomi Corp 47.9% 92/192 2
51 Claude Sonnet 4.5 20250929 High 32k no overall score anthropic 46.4% 89/192 2
52 Gemini 3 Flash (Preview) no overall score google 45.3% 87/192 2
53 Z-Ai/GLM 4.7 no overall score Z.ai (Zhipu AI) 44.8% 86/192 2
54 Claude Sonnet 4.5 Anthropic 42.7% 82/192 2 31
55 GPT-5.4 mini High no overall score openai 42.7% 82/192 2
56 MiMo V2.5 no overall score Xiaomi Corp 42.2% 81/192 2
57 Deepseek V4 Flash High Preview no overall score deepseek 41.7% 80/192 2
58 Inkling no overall score Thinking Machines 41.7% 80/192 2
59 Claude 4.1 Opus Anthropic 40.6% 78/192 2 32
60 Qwen3.5 397B A17B no overall score Alibaba 40.1% 77/192 2
61 OpenAI/GPT-5.2 no overall score openai 39.8% 76.5/192 2
62 Gemma 4 31B no overall score google 39.1% 75/192 2
63 GLM-5V-Turbo no overall score zai 36.5% 70/192 2
64 Kimi K2.5 Instant no overall score moonshot 36.5% 70/192 2
65 Grok 4.1 Thinking no overall score xai 33.9% 65/192 2
66 Grok 4.3 no overall score xAI 32.3% 62/192 2
67 Gemma 4 26B A4B no overall score Google DeepMind 31.3% 60/192 2
68 MiniMax-M2.7 no overall score MiniMax 29.2% 56/192 2
69 GPT-5.1 no overall score openai 28.6% 55/192 2
70 Inkling Small no overall score Thinking Machines 28.1% 54/192 2
71 Muse Glimmer no overall score meta 25.0% 48/192 2
72 DeepSeek V3.2 Thinking no overall score deepseek 24.0% 46/192 2
73 Gemini 2.5 Pro Google DeepMind 24.0% 46/192 2 34
74 Qwen3.5 122B A10B no overall score alibaba 21.9% 42/192 2
75 Kimi K2 Thinking Turbo no overall score Moonshot 21.4% 41/192 2
76 Z-AI/GLM 4.6 no overall score Z.ai (Zhipu AI),Tsinghua University 20.8% 40/192 2
77 MiniMax-M2.5 no overall score MiniMax 20.8% 40/192 2
78 Deepseek/DeepSeek-V3.2 no overall score DeepSeek 20.3% 39/192 2
79 Minimax M2.1 Preview no overall score minimax 20.3% 39/192 2
80 Gemini 3.1 Flash Lite Preview no overall score google 19.8% 38/192 2
81 Qwen3.5 27B no overall score Alibaba 18.8% 36/192 2
82 Hunyuan Hy3 Preview no overall score tencent 18.2% 35/192 2
83 Mistral Medium 3.5 no overall score mistral 18.2% 35/192 2
84 X-Ai/Grok 4.1 Fast Reasoning no overall score xAI 17.7% 34/192 2
85 Claude Haiku 4.5 no overall score Anthropic 17.2% 33/192 2
86 Solar Pro 4 no overall score upstage 17.2% 33/192 2
87 DeepSeek/DeepSeek-V3.2-Exp no overall score DeepSeek 16.1% 31/192 2
88 Mistral Large 3 no overall score mistral 10.9% 21/192 2
89 Mimo-V2-Flash no overall score Xiaomi Corp 10.4% 20/192 2
90 Qwen3 Coder 480B A35B Instruct no overall score Alibaba 10.4% 20/192 2
91 Qwen3.5 35B A3B no overall score alibaba 9.4% 18/192 2
92 Minimax/Minimax-M2 no overall score minimax 8.3% 16/192 2
93 Qwen3.5 Flash no overall score Alibaba 8.3% 16/192 2
94 X-Ai/Grok-4-Fast-Reasoning no overall score xai 5.7% 11/192 2
95 Trinity Large Thinking no overall score 5.2% 10/192 2
96 Mercury 2 no overall score Inception Labs 1.6% 3/192 2
97 Granite 4.1 8B no overall score ibm 1.0% 2/192 2
Measured, not ranked

Where the rest of the field is

142measured, cannot be scored

37 models hold a Ferrox Index. 142 more are measured on at least one category and cannot be scored. A model needs 2 independent benchmarks in EACH of agents, coding, reasoning before it can hold a Ferrox Index. Models below that bar are listed with exactly what they are short of. Almost none of them are unmeasured; they are measured unevenly, and the gap is in the public record, not in our reading of it.

The bottleneck is not close: coding is missing for 125, agents is missing for 125, reasoning is missing for 11 of them. Agent and coding evaluations are run on far fewer models than reasoning evaluations are, and that gap is in the public record rather than in our reading of it. A model here needs one more independently run benchmark in the category shown, and it is comparable.

ModelOrganisationBenchmarks heldQualifies inShort of
GPT-4o (2024-11-20) OpenAI 10 agents, reasoning coding 1/2
Claude Haiku 4.5 Anthropic 9 agents, reasoning coding 1/2
GPT-4.1 (2025-04-14) OpenAI 9 coding, reasoning agents 0/2
Grok 4.0709 xAI 9 agents, reasoning coding 1/2
o4 mini (2025-04-16) OpenAI 9 coding, reasoning agents 1/2
Z-Ai/GLM 5 Z.ai (Zhipu AI) 9 coding, reasoning agents 1/2
DeepSeek V4 Pro DeepSeek 8 coding, reasoning agents 1/2
Gemini 3.5 Flash Google 8 coding, reasoning agents 1/2
GPT-5 mini (2025-08-07) OpenAI 8 coding, reasoning agents 1/2
Inkling open Thinking Machines 8 coding, reasoning agents 1/2
o1 (2024-12-17) OpenAI 8 agents, reasoning coding 0/2
DeepSeek-R1-0528 DeepSeek 7 coding, reasoning agents 1/2
Gemini 3.6 Flash Google DeepMind 7 coding, reasoning agents 1/2
Gemini 3.7 Flash Google DeepMind 7 coding, reasoning agents 1/2
GPT-5.4 mini (2026-03-17) OpenAI 7 coding, reasoning agents 1/2
Qwen3.6 Plus Alibaba 7 coding, reasoning agents 0/2
Claude 3.5 Sonnet (2024-06-20) Anthropic 6 agents, reasoning coding 0/2
Claude 3 Opus (2024-02-29) Anthropic 5 agents, reasoning coding 0/2
DeepSeek V4 Flash 0731 open DeepSeek 5 coding, reasoning agents 0/2
GPT-5.3 Codex OpenAI 5 agents, coding reasoning 0/2
Qwen3.7 Max Alibaba 5 coding, reasoning agents 1/2
Nemotron 3 Ultra (free) NVIDIA 4 agents, reasoning coding 0/2
o1 Preview (2024-09-12) OpenAI 4 agents, reasoning coding 0/2
Claude 3.5 Sonnet (2024-10-22) Anthropic 7 reasoning agents 1/2, coding 0/2
Gemini 3.5 Flash Lite Google DeepMind 7 reasoning agents 1/2, coding 1/2
GPT-4.5 Preview (2025-02-27) OpenAI 7 reasoning agents 1/2, coding 1/2
Z-Ai/GLM 4.7 Z.ai (Zhipu AI) 7 reasoning agents 1/2, coding 1/2
DeepSeek-R1 DeepSeek 6 reasoning agents 1/2, coding 0/2
DeepSeek-V3 DeepSeek 6 reasoning agents 1/2, coding 1/2
Gemini 2.5 Pro Preview 0605 Google DeepMind 6 reasoning agents 1/2, coding 1/2
GPT-4.1 Nano (2025-04-14) OpenAI 6 reasoning agents 0/2, coding 1/2
GPT-5 Nano (2025-08-07) OpenAI 6 reasoning agents 1/2, coding 1/2
GPT-5.4 Nano (2026-03-17) OpenAI 6 reasoning agents 1/2, coding 1/2
Grok 4.3 xAI 6 reasoning agents 1/2, coding 1/2
Inkling Small open Thinking Machines 6 reasoning agents 1/2, coding 0/2
MiniMax-M2.5 MiniMax 6 agents coding 1/2, reasoning 1/2
Muse Spark 1.1 Meta AI 6 agents coding 1/2, reasoning 0/2
o1 mini (2024-09-12) OpenAI 6 reasoning agents 1/2, coding 0/2
o3 mini (2025-01-31) OpenAI 6 reasoning agents 1/2, coding 0/2
OpenAI GPT-4.1 Mini OpenAI 6 reasoning agents 0/2, coding 1/2
Qwen3.5 Flash Alibaba 6 reasoning agents 0/2, coding 1/2
Qwen3.6 Max Preview Alibaba 6 reasoning agents 0/2, coding 1/2
Qwen3.8 Max Alibaba 6 reasoning agents 1/2, coding 1/2
Z-AI/GLM 4.6 Z.ai (Zhipu AI),Tsinghua University 6 agents coding 1/2, reasoning 1/2
Claude 3.5 Haiku (2024-10-22) Anthropic 5 reasoning agents 0/2, coding 0/2
DeepSeek-V3-0324 DeepSeek 5 reasoning agents 1/2, coding 1/2
Gemini 1.5 Pro 002 Google DeepMind 5 reasoning agents 0/2, coding 0/2
Gemini 2.0 Flash 001 Google DeepMind,Google 5 reasoning agents 0/2, coding 0/2
Gemini 3.1 Flash Lite Google 5 reasoning agents 1/2, coding 1/2
GPT-4o (2024-08-06) OpenAI 5 reasoning agents 0/2, coding 0/2
GPT-4o-mini (2024-07-18) OpenAI 5 reasoning agents 0/2, coding 0/2
GPT-OSS 20B OpenAI 5 reasoning agents 1/2, coding 1/2
Grok 3 mini Beta xAI 5 reasoning agents 0/2, coding 1/2
Kimi K2 Thinking Moonshot 5 agents coding 1/2, reasoning 1/2
Llama 4 Scout 17B 16E Instruct open Meta AI 5 reasoning agents 0/2, coding 0/2
MiniMax-M2.7 MiniMax 5 agents coding 1/2, reasoning 0/2
Mistral Medium 3 Mistral AI 5 reasoning agents 0/2, coding 0/2
Muse Spark Meta AI 5 reasoning agents 0/2, coding 0/2
Qwen3.5 Plus Alibaba 5 reasoning agents 1/2, coding 1/2
Qwen3.7 Plus Alibaba 5 reasoning agents 1/2, coding 1/2
DeepSeek V4 Pro 0813 DeepSeek 4 reasoning agents 0/2, coding 1/2
Gemini 1.5 Flash 002 Google DeepMind 4 reasoning agents 0/2, coding 0/2
Gemini 2.0 Flash Thinking Exp 01.21 Google DeepMind,Google 4 reasoning agents 0/2, coding 0/2
Gemini 2.5 Flash 0520 Google DeepMind 4 reasoning agents 0/2, coding 1/2
Gemma 3.27B It Google DeepMind 4 reasoning agents 0/2, coding 1/2
GPT-4 Turbo (2024-04-09) OpenAI 4 reasoning agents 1/2, coding 0/2
GPT-4.0314 OpenAI 4 reasoning agents 1/2, coding 0/2
GPT-5.1 Codex OpenAI 4 agents coding 1/2, reasoning 0/2
GPT-5.4 Pro (2026-03-05) OpenAI 4 reasoning agents 0/2, coding 0/2
GPT-5.5 Pre Release OpenAI 4 reasoning agents 0/2, coding 1/2
Grok 3 Beta xAI 4 reasoning agents 0/2, coding 1/2
Kimi K2 Thinking Turbo Moonshot 4 reasoning agents 0/2, coding 0/2
Llama 4 Maverick 17b 128e Instruct open Meta AI 4 reasoning agents 0/2, coding 1/2
Llama 4 Maverick 17B FP8 open Meta AI 4 reasoning agents 0/2, coding 1/2
Mistral Large 2411 Mistral AI 4 reasoning agents 0/2, coding 0/2
Qwen2.5 72B Instruct Alibaba 4 reasoning agents 1/2, coding 0/2
Qwen2.5-Max-2025-01-25 Alibaba 4 reasoning agents 0/2, coding 0/2
Qwen3 235B A22B Thinking 2507 Alibaba 4 reasoning agents 0/2, coding 0/2
Qwen3 Max (2025-09-23) Alibaba 4 reasoning agents 0/2, coding 1/2
Qwen3.5 397B A17B Alibaba 4 reasoning agents 0/2, coding 0/2
Qwen3.6 Flash Alibaba 4 reasoning agents 0/2, coding 1/2
Qwq 32B open Alibaba 4 reasoning agents 0/2, coding 1/2
Anthropic: Claude 3 Haiku Anthropic 3 reasoning agents 0/2, coding 0/2
Claude 3 Sonnet (2024-02-29) Anthropic 3 reasoning agents 0/2, coding 0/2
DeepSeek Chat DeepSeek 3 reasoning agents 0/2, coding 1/2
Deepseek Llm 67B Chat DeepSeek 3 reasoning agents 0/2, coding 0/2
DeepSeek Reasoner DeepSeek 3 reasoning agents 0/2, coding 1/2
Gemini 1.5 Flash 001 Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemini 1.5 Flash 8B 001 Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemini 1.5 Pro 001 Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemini 2.5 Flash Preview Google DeepMind 3 reasoning agents 0/2, coding 1/2
Gemini 2.5 Pro Experimental 0325 Google DeepMind 3 reasoning agents 0/2, coding 1/2
Gemini 2.5 Pro Preview 0506 Google DeepMind 3 reasoning agents 0/2, coding 1/2
Gemma 2 27B Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemma 2.9B It Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemma 3 12B IT Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemma 3 4B IT Google DeepMind 3 reasoning agents 0/2, coding 0/2
Gemma 4 31B IT Google DeepMind 3 reasoning agents 0/2, coding 1/2
GLM-5.3-Flash Z.ai (Zhipu AI) 3 reasoning agents 0/2, coding 0/2
GPT 5.1 Codex Max OpenAI 3 agents coding 1/2, reasoning 0/2
GPT-3.5 Turbo 0125 OpenAI 3 reasoning agents 0/2, coding 0/2
GPT-4.0613 OpenAI 3 reasoning agents 0/2, coding 0/2
GPT-4o (2024-05-13) OpenAI 3 reasoning agents 0/2, coding 0/2
GPT-5.5 Instant OpenAI 3 reasoning agents 0/2, coding 0/2
GPT-5.5 Pro Pre Release OpenAI 3 reasoning agents 0/2, coding 0/2
Grok 2.1212 xAI 3 reasoning agents 0/2, coding 0/2
Kimi K2p5 Moonshot 3 reasoning agents 0/2, coding 0/2
Llama 3.1 405B Instruct open Meta AI 3 reasoning agents 1/2, coding 0/2
Llama 3.1 70b Instruct Meta AI 3 reasoning agents 0/2, coding 0/2
Llama 3.1 8B Instruct Meta AI 3 reasoning agents 0/2, coding 0/2
Llama 3.3 70B Instruct Meta AI 3 reasoning agents 0/2, coding 0/2
Magistral Small 2506 Mistral AI 3 reasoning agents 0/2, coding 0/2
Meta Llama 3.70B Instruct Meta AI 3 reasoning agents 1/2, coding 0/2
Mistral Large 2402 Mistral AI 3 reasoning agents 0/2, coding 0/2
Mistral Large 2407 Mistral AI 3 reasoning agents 0/2, coding 0/2
Phi 4 Microsoft Research 3 reasoning agents 0/2, coding 0/2
Qwen3.5 9B Alibaba 3 reasoning agents 1/2, coding 0/2
Qwen3.6 35B A3B Alibaba 3 reasoning agents 1/2, coding 0/2
Claude 2.0 Anthropic 2 reasoning agents 0/2, coding 0/2
Claude 2.1 Anthropic 2 reasoning agents 0/2, coding 0/2
Gemini 1.0 Pro 001 Google DeepMind 2 reasoning agents 0/2, coding 0/2
Gemma 3.1B It Google DeepMind 2 reasoning agents 0/2, coding 0/2
Gemma 4 26B A4B IT Google DeepMind 2 reasoning agents 0/2, coding 0/2
GPT-5 Codex OpenAI 2 agents coding 0/2, reasoning 0/2
GPT-5 Pro (2025-10-06) OpenAI 2 reasoning agents 0/2, coding 0/2
Grok 4.20 (Reasoning) xAI 2 reasoning agents 0/2, coding 0/2
Hermes 2 Theta Llama 3.70B Nous Research,Arcee AI 2 reasoning agents 0/2, coding 0/2
Kimi K2 0905 Moonshot 2 coding agents 0/2, reasoning 0/2
Llama 2.70B Chat Hf Meta AI 2 reasoning agents 0/2, coding 0/2
Llama 3.1 Tulu 3.70B Dpo Allen Institute for AI,University of Washington 2 reasoning agents 0/2, coding 0/2
Llama-3.2-90B-Vision-Instruct open Meta AI 2 reasoning agents 0/2, coding 0/2
Meta Llama 3.8B Instruct Meta AI 2 reasoning agents 0/2, coding 0/2
Mistral Medium 3.5 open Mistral AI 2 coding agents 0/2, reasoning 0/2
Mistral Small 2501 Mistral AI 2 reasoning agents 0/2, coding 0/2
Mistral Small 2503 Mistral AI 2 reasoning agents 0/2, coding 0/2
Qwen 3 8B Alibaba 2 reasoning agents 0/2, coding 0/2
Qwen Plus (2025-01-25) Alibaba 2 reasoning agents 0/2, coding 0/2
Qwen Turbo (2024-11-01) Alibaba 2 reasoning agents 0/2, coding 0/2
Qwen2.5 32B Instruct open Alibaba 2 reasoning agents 0/2, coding 0/2
Qwen3.6 27B Alibaba 2 reasoning agents 0/2, coding 0/2
Qwen3.7 Flash Alibaba 2 reasoning agents 0/2, coding 0/2
R1 Distill Llama 70B DeepSeek 2 reasoning agents 0/2, coding 0/2

All 142, closest first. A highlighted row is one category away from the board. Models with no qualifying category at all are not listed: there is nothing measured about them to report.

What moves a score

The score you were shown was chosen

32/37move on which run gets quoted

The same model, on the same benchmark, run under a different harness or a different reasoning effort, produces a different score. Whoever publishes gets to pick from that range. Each rail below is the full range one model actually produced, with the value we publish marked inside it and the best configuration marked at the top. We take the median. A best-of rule takes the top mark.

On 32 of the 37 models here, the number moves more than 5 points depending on which run gets quoted. GPT-5.6 Luna spans 65.6 points on DeepSWE across 5 configurations of itself.

range across configurations what we publish (median) best configuration
GPT-5.6 Luna DeepSWE · 5 configs
65.6
GPT-5.6 Terra ARC-AGI-2 · 6 configs
65.2
GLM-5.2 OTIS Mock AIME 2024-2025 · 3 configs
57.5
GPT-5.2 (2025-12-11) ARC-AGI-2 · 7 configs
52.1
GPT-5.5 ARC-AGI-2 · 5 configs
51.7
GPT-5.1 (2025-11-13) OTIS Mock AIME 2024-2025 · 4 configs
50.8
GPT-5.6 Sol ARC-AGI-2 · 6 configs
50.0
Kimi K3 ARC-AGI-2 · 3 configs
48.1
GPT-5.4 (2026-03-05) ARC-AGI-2 · 5 configs
44.8
GPT-5 (2025-08-07) OTIS Mock AIME 2024-2025 · 3 configs
44.7
MiniMax-M3 OTIS Mock AIME 2024-2025 · 2 configs
44.4
Claude Sonnet 4.5 OTIS Mock AIME 2024-2025 · 4 configs
42.2
050100

Widest benchmark per model, 12 shown of 32 that move more than 5 points. Adjacent ranks on this board are separated by as little as 0.215 index points, so a five point swing on one benchmark reorders the table. Scores are normalised to 0 to 100 for comparison; the published values are on each benchmark panel below.

Every benchmark, on its own terms

32benchmarks · 31% of the grid measured

One panel per benchmark, best first, the published value printed on the bar. The leader is in orange and the rest of the field is grey. Under each panel: who ran it, under what licence, and how many ranked models are missing from it. That last line is the one no vendor chart carries.

Behind every bar there is usually more than one number. 321 of these model-and-benchmark cells hold two or more published measurements, and 14 of those hold measurements that disagree while the record says the configuration was the same. The widest gap on the board is 65.6 points, on one model at one runner, moved by nothing but the effort dial. Every claim we can check →

Terminal-Bench 2

agents
14.2% scaled to this field 82.2%
and 5 more measured
Terminal-Bench CC-BY-4.0
20 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness

BrowseComp

agents
59.4%
various; recorded per row facts, not expression
36 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

IMOAnswerBench

reasoning
86.8%
various; recorded per row facts, not expression
36 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

SWE-bench Multilingual

coding
74.8%
various; recorded per row facts, not expression
36 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

AIME

reasoning
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

CyberGym

agents
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

HMMT

reasoning
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

MCP-Atlas

agents
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

SWE-bench Pro

coding
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

Tau2-Bench Airline

agents
no ranked model has been measured on this benchmark
Tau2-Bench (Sierra) MIT
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

Tau2-Bench Banking

agents
no ranked model has been measured on this benchmark
Tau2-Bench (Sierra) MIT
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

Tau2-Bench Retail

agents
no ranked model has been measured on this benchmark
Tau2-Bench (Sierra) MIT
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

Tau2-Bench Telecom

agents
no ranked model has been measured on this benchmark
Tau2-Bench (Sierra) MIT
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

Tool-Decathlon

agents
no ranked model has been measured on this benchmark
various; recorded per row facts, not expression
37 of 37 ranked models have no measurement here
scores are harness-keyed: the same model scores differently under a different harness
this evaluator publishes no uncertainty

The holes

coverage

31% of the grid is measured. 812 of 1,184 model-benchmark pairs have no published number from anyone.

  • 23 of 32 benchmarks cover fewer than half the ranked field
  • 19 publish no uncertainty at all
  • 19 are harness-keyed, so the harness is part of the result
  • 104 vendor self-reports held out of every score
A missing cell is not a zero and is never treated as one. It is the reason this board exists.
Method

How this is built, and what it refuses

82.4%agreement with evidence the ordering never saw

Category ordering is head to head. ordered by Bradley-Terry ratings fitted to every pairwise comparison between two models on the benchmarks they BOTH sat. This orders the category boards, and since 2026-08-29 it orders the Ferrox Index too: a category contributes its position in this fit, weighted by the rubric, never the raw composite it used to contribute. the figure shown is the observed win share: of every head-to-head comparison a model actually played, the share it won. The fitted rating orders the board; it is not published as a score, because under near-perfect transitivity its magnitude reflects the prior rather than a measurement. We rank this way because placement removed the difficulty of a benchmark but not the strength of the field that sat it: the models measured on Terminal-Bench 4.0 averaged overall rank 11.6, while Aider Polyglot averaged 26.8, so sitting an easy room was worth real position. And not on direct results alone because averaging direct results alone compares models over different opponent sets, and 107 of 561 model pairs in the agents category share no benchmark at all. A fitted rating ranks A above C on the strength of both having played B.

every model also plays one drawn game against an average opponent, so a model that won everything has a finite rating and a model with little evidence is pulled gently toward the middle. ratings are normalised so the geometric mean is 1, which pins the scale Bradley-Terry leaves free and makes "an average opponent" mean exactly that. The published figure is the chance of winning one comparison against such an opponent. the comparison graph must be one connected component, which it is in every category on this snapshot, with no isolated model. the fit runs to convergence rather than to a cap, reaching it in 544 to 1833 iterations depending on the category, and the ordering it reaches is identical to one left running 200,000 iterations, in every category. It also barely moves with the strength of the prior: rank correlation against an unregularised fit stays at or above 0.972 from prior 2 to prior 20, four times either side of the value used. Both are asserted by tests rather than assumed. What it does not fix: a benchmark is not a fair coin. Two models separated by a hair and two separated by a mile both count as one win, so this measures who beats whom and not by how much. The margins are on the benchmark panels. placement changes what a category board SHOWS as evidence and never orders anything; the ordering is head to head. Since 2026-08-29 the Ferrox Index is built from those head-to-head positions rather than from the raw category scores, so it is no longer true that the index is untouched by any of this.

How often this board agrees with evidence it never saw: 82.4%. Hide one correlation family of benchmarks. Refit every category board without it, re-derive each category position, re-run the index and re-rank. Then ask the hidden family which of two models is better, and count how often the ordering had already put that one first. Repeat for every family and pool the pairs. An ordering can only be said to have predicted something if it did not already know the answer, so this is the number in the corner and the one to judge the board by. On this snapshot it is measured over 17 folds and 2311 held-out pairs. An arbitrary ordering of the same field scores 51.3% on the same pairs, which is what chance looks like and what this has to be read against. A held-out pair is decided by less evidence than a published pair, often by a single family, so its verdict is noisier and the reachable ceiling is below 100 by an amount this does not estimate. Read it against the arbitrary-ordering baseline, not against 100.

Refitting the board without a family moves it a little and never reorders it wholesale: mean Spearman 0.9992 against the published order across the 17 folds, worst 0.9967. 3 of those folds decided no pair between two ranked models and so validated nothing; they are counted here rather than quietly averaged in as agreement.

And how often it agrees with the evidence it was built from: 95.3%. For every pair of ranked models, count the benchmarks both of them sat and see which won more of them. Pairs that split evenly, or share no benchmark, are set aside as undecided. The agreement rate is the share of the decided pairs this ordering puts the right way round. On this snapshot 619 of 666 pairs are decided that way, and the ordering above puts 29 of them the wrong way round. Those 29 are listed below rather than counted, because a number you cannot check is worth nothing. This figure is measured on the same comparisons the ordering was fitted to, so it is a statement about internal consistency and not about predictive accuracy. It was published here on its own until 2026-08-30, in a position that read as validation, and it should not have been. It counts benchmarks won, never by how much, so a hair and a landslide weigh the same. It cannot reach 100% on evidence that is intransitive, and it has no opinion about which categories matter. It grades orderings against each other; it does not certify one as correct.

The gap between the two is not all overfitting, and reading it that way would be as wrong as the original claim. Graded on one hidden family instead of on everything a pair shares, the published ordering itself scores 83.3%: that drop is the price of thinner evidence per pair, and it is paid whether or not the ordering saw the fold. Taking the fold out of the fit costs a further 0.9 points, on identical pairs with nothing else varied. That last number is the overfitting, and it is the one worth arguing about.

Ranked aboveBeaten byShared benchmarks
GPT-OSS 120B Claude 4 Sonnet lost 1 to 5 of 7
GPT-5.1 (2025-11-13) GPT-5 (2025-08-07) lost 2 to 6 of 9
Claude 4 Opus Gemini 2.5 Pro lost 1 to 5 of 6
GPT-5.4 (2026-03-05) GPT-5.5 lost 2 to 5 of 7
Claude Opus 4.6 Gemini 3.1 Pro Preview lost 4 to 7 of 11
GPT-5.6 Sol Claude Fable 5 lost 3 to 6 of 9
Claude 4.5 Opus GLM 5.1 lost 3 to 6 of 9
GPT-OSS 120B Claude 3.7 Sonnet lost 1 to 4 of 5
Claude 3.7 Sonnet Claude 4 Sonnet lost 3 to 6 of 9
GLM-5.2 Gemini 3 Pro Preview lost 2 to 4 of 6

Showing the 10 widest of 29. An ordering cannot satisfy every pair when the evidence itself is intransitive, so this number will never reach 100 and a board claiming it did would be telling you something else.

A grade is a position, never an absolute. The field on this snapshot runs from 99.5 down to 23.1, so grading against a 0 to 100 scale would put the best model in the world in the middle of the alphabet. a grade boundary never splits a pair the evidence cannot separate; the lower row is lifted, capped at one band, and the lift is named on the row.

GradePercentile floorModelsRanksMeaning
A+ 904 1 to 4top of the measured field
A 758 5 to 12front rank
B 555 13 to 17above the field median
C 357 18 to 24below the field median
D 157 25 to 31back of the field
E 06 32 to 37bottom of the measured field

grades are for legibility. They never enter the Ferrox Index and never override a tie band.

SNAPSHOT2026-08-30-62e4e85f043b
SOURCE MANIFESTdc46131315046f1e
AS OF2026-08-30
RUBRICv1.1.0 UNCALIBRATED

Measured by