THE FORGE RANKINGS·2026-08-30

GLM 5.1

Z.ai (Zhipu AI)
25
RANK of 37 ranked
FERROX INDEX58.7 weighted mean of category positions
GRADED back of the field
EVIDENCE13 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

13of 37 measured
44.2
13.1median 38.963.0

Coding

20of 37 measured
49.7
22.4median 50.070.7

Reasoning

19of 37 measured
62.8
18.1median 62.893.0

Preference

7of 15 measured
72.3
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

OTIS Mock AIME 2024-2025

12/36
93.3% best 99.7% · Claude Fable 5
Epoch AI CC-BY-4.0

ALE-Bench

26/35
887 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0

LMArena Agent

19/21
-0.008 best 0.127 · Claude Opus 5
LMArena CC-BY-4.0

FrontierMath

10/21
33.4% best 47.6% · GPT-5.4 (2026-03-…
Epoch AI CC-BY-4.0

LMArena Text (style-controlled)

10/20
1466 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

LMArena WebDev

7/16
1509 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

SWE-bench Multilingual

1/1
74.8% best 74.8% · GLM 5.1
NVIDIA facts, not expression NeMo Evaluator SDK model-card defaults

Not measured

19/32
  • ARC-AGI-2
  • APEX Agents
  • DeepSWE
  • FrontierCode
  • METR time horizons
  • Terminal-Bench 4.0
  • Aider Polyglot
  • OSWorld
  • Cybench
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
BrowseComp agents 59.4% not published NVIDIA facts, not expression Tavily plus terminal workspace / model-card defaults
LMArena Agent agents -0.008 11.2 to 16.9 LMArena CC-BY-4.0 source
Terminal-Bench 2 agents 59.3% not published NVIDIA CC-BY-4.0 NeMo Evaluator SDK; Harbor / model-card defaults source
ALE-Bench coding 887 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 source
SWE-bench Multilingual coding 74.8% not published NVIDIA facts, not expression NeMo Evaluator SDK / model-card defaults
SWE-bench Verified (Epoch's own run) coding 73.2% not published vLLM-Ascend project CC-BY-4.0 GPU baseline; 10 groups over 3 rounds / not disclosed source
LMArena Text (style-controlled) preference 1466 80.3 to 81.5 LMArena CC-BY-4.0 source
LMArena WebDev preference 1509 62.7 to 64.5 LMArena CC-BY-4.0 source
FrontierMath reasoning 33.4% 28.0 to 38.9 Epoch AI CC-BY-4.0 source
GPQA Diamond (Epoch's own run) reasoning 89.9% 85.7 to 94.1 Epoch AI CC-BY-4.0 source
Humanity's Last Exam reasoning 27.2% not published NVIDIA CC-BY-4.0 NeMo Evaluator SDK / model-card defaults source
IMOAnswerBench reasoning 86.8% not published NVIDIA facts, not expression NeMo Evaluator SDK / model-card defaults
OTIS Mock AIME 2024-2025 reasoning 93.3% 86.0 to 100.0 Epoch AI CC-BY-4.0 source

Configurations rolled up: 7. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: ascend-a3-dual-node, ascend-a5-w4a4c8, gpu-baseline;-10-groups-over-3-rounds, nemo-evaluator-sdk, nemo-evaluator-sdk;-harbor, tavily-plus-terminal-workspace.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.