All 1 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.
Matched on which variant of the benchmark was sat, and on nothing else: the vendor ran it on their harness and the independent party on theirs, and that difference is the comparison rather than an obstacle to it. A vendor number with no independent run of the same variant carries no delta, because a number that cannot be compared honestly must not be compared.
| Variant | Vendor claims | Independently measured | Difference |
|---|---|---|---|
| 2.1 | 62.0 | 59.3 from 1 run | +2.7 |
Published by the model's own maker, on a variant of this benchmark that no independent party we hold has run. They are recorded and they are shown. They can never move a rank.
| Variant | Claimed |
|---|---|
| 2.0 best self-reported Claude Code | 69.0 |
| 2.0 Terminus-2 | 63.5 |
4 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.
| Score | Provenance | Configuration | Who ran it | Where to check it |
|---|---|---|---|---|
| 69.0 / 100
as published: 69.0% |
CLAIMED
shown, never scored |
glm-5.1 harness Claude Code effort not disclosed variant 2.0 best self-reported Claude Code |
Z.AI via Model benchmark dossier facts, not expression |
checked: this figure is not on the page today2026-08-30 huggingface.coGLM-5.1 |
| 63.5 / 100
as published: 63.5% |
CLAIMED
shown, never scored |
glm-5.1 harness model-card evaluation effort not disclosed variant 2.0 Terminus-2 |
Z.AI via Model benchmark dossier facts, not expression |
checked: this figure is not on the page today2026-08-30 huggingface.coGLM-5.1 |
| 62.0 / 100
as published: 62.0% |
CLAIMED
shown, never scored |
glm-5.1 harness later GLM-5.2 comparison effort not disclosed variant 2.1 |
Z.AI via Model benchmark dossier facts, not expression |
checked: found, but not in a block naming only this model2026-08-30 github.comOn standard coding benchmarks, GLM-5.2 is the strongest open-source model, improving on GLM-5.1 by a wide margin: 81.0 vs. 62.0 on Terminal-Bench 2.1 and 62.1 v |
| 59.3 / 100
as published: 59.3% |
INDEPENDENT
counts toward the ranking |
glm-5.1 harness NeMo Evaluator SDK; Harbor effort model-card defaults variant 2.1 |
NVIDIA via Model benchmark dossier facts, not expression |
not checked: the page is a format we cannot read2026-08-30 research.nvidia.com |
Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.