3 independent runners have measured this, which is rare: on this snapshot only 12 of 321 claim pages can say it.
5 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.
| Score | Provenance | Configuration | Who ran it | Where to check it |
|---|---|---|---|---|
| 76.2 / 100
as published: 76.2% |
INDEPENDENT
counts toward the ranking |
glm-5.1 harness NeMo Evaluator SDK effort model-card defaults variant verified |
NVIDIA via Model benchmark dossier facts, not expression |
not checked: the page is a format we cannot read2026-08-30 research.nvidia.com |
| 74.2 / 100
as published: 74.2% |
INDEPENDENT
counts toward the ranking |
glm-5.1 | Epoch AI 2026-05-15T10:07:59.494Z CC-BY-4.0 |
NO_URL2026-08-30 benchmark leaderboardthe leaderboard for this benchmark, not a link to this row |
| 73.2 / 100
as published: 73.2% |
INDEPENDENT
counts toward the ranking |
glm-5.1 harness GPU baseline; 10 groups over 3 rounds effort not disclosed variant verified |
vLLM-Ascend project via Model benchmark dossier 2026-08-15T00:00:00.000Z facts, not expression |
checked: found, but not in a block naming only this model2026-08-30 github.com"@context":"https://schema.org","@type":"DiscussionForumPosting","headline":"[DeepSeek-V4-Flash/GLM-5.1] Intermittent SWE benchmark failures and score variabili |
| 72.0 / 100
as published: 72.0% |
INDEPENDENT
counts toward the ranking |
glm-5.1 harness Ascend A5 W4A4C8 effort not disclosed variant verified |
vLLM-Ascend project via Model benchmark dossier 2026-08-15T00:00:00.000Z facts, not expression |
checked: found, but not in a block naming only this model2026-08-30 github.com"@context":"https://schema.org","@type":"DiscussionForumPosting","headline":"[DeepSeek-V4-Flash/GLM-5.1] Intermittent SWE benchmark failures and score variabili |
| 71.4 / 100
as published: 71.4% |
INDEPENDENT
counts toward the ranking |
glm-5.1 harness Ascend A3 dual-node effort not disclosed variant verified |
vLLM-Ascend project via Model benchmark dossier 2026-08-15T00:00:00.000Z facts, not expression |
checked: this figure appears in the page’s row for this model2026-08-30 github.comGLM-5.1: A3 dual-node: 71.4% |
Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.