All 0 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.
Published by the model's own maker, on a variant of this benchmark that no independent party we hold has run. They are recorded and they are shown. They can never move a rank.
| Variant | Claimed |
|---|---|
| 2.0 Terminus-2 | 56.2 |
| 2.0 Claude Code | 56.2 |
| no variant stated | 52.4 |
| 2.0 Terminus-2 official | 52.4 |
4 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.
| Score | Provenance | Configuration | Who ran it | Where to check it |
|---|---|---|---|---|
| 56.2 / 100
as published: 56.2% |
CLAIMED
shown, never scored |
harness Terminus-2 effort not disclosed variant 2.0 Terminus-2 |
Z.AI via Model benchmark dossier facts, not expression |
checked: this figure is not on the page today2026-08-30 huggingface.coGLM-5 |
| 56.2 / 100
as published: 56.2% |
CLAIMED
shown, never scored |
harness Claude Code 2.1.14; five runs effort not disclosed variant 2.0 Claude Code |
Z.AI via Model benchmark dossier facts, not expression |
checked: this figure is not on the page today2026-08-30 huggingface.coGLM-5 |
| 52.4 / 100
as published: 52.4% |
UNATTRIBUTED
shown, never scored |
harness Terminus 2 | Terminal-Bench via Epoch AI 2026-02-23 CC-BY-4.0 |
NO_URL2026-08-30 benchmark leaderboardthe leaderboard for this benchmark, not a link to this row |
| 52.4 / 100
as published: 52.4% |
CLAIMED
shown, never scored |
harness evaluator historical row unavailable effort not disclosed variant 2.0 Terminus-2 official |
Z.AI model registry via Model benchmark dossier facts, not expression |
checked: this figure is not on the page today2026-08-30 huggingface.coGLM-5 |
Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.