Claim · coding

GLM 5.1on SWE-bench Verified (Epoch's own run)

71.4 to 76.2 the range across every scoreable measurement we hold
4.8 points between the highest and the lowest
73.2 the number the board uses: the lower median, a real row with its own provenance
5 declared configurations, from 3 runners
1 / 4 rows whose cited page still shows this figure, refetched and checked by us

3 independent runners have measured this, which is rare: on this snapshot only 12 of 321 claim pages can say it.

2 of 4 cited pages no longer show the figure we publish. The page was refetched and split at its own structural boundaries, and no block that names this model carries this number. That does not make the measurement wrong: a leaderboard is a living page and a figure recorded when it was there can be gone by the time you follow the link. It does mean the citation, as printed, no longer supports the number, and this is the only place we know of that says so about its own sources.

Every measurement

5 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
76.2 / 100
as published: 76.2%
INDEPENDENT
counts toward the ranking
glm-5.1
harness NeMo Evaluator SDK
effort model-card defaults
variant verified
NVIDIA
via Model benchmark dossier
facts, not expression
not checked: the page is a format we cannot read2026-08-30
research.nvidia.com
74.2 / 100
as published: 74.2%
INDEPENDENT
counts toward the ranking
glm-5.1 Epoch AI
2026-05-15T10:07:59.494Z
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
73.2 / 100
as published: 73.2%
INDEPENDENT
counts toward the ranking
glm-5.1
harness GPU baseline; 10 groups over 3 rounds
effort not disclosed
variant verified
vLLM-Ascend project
via Model benchmark dossier
2026-08-15T00:00:00.000Z
facts, not expression
checked: found, but not in a block naming only this model2026-08-30"@context":"https://schema.org","@type":"DiscussionForumPosting","headline":"[DeepSeek-V4-Flash/GLM-5.1] Intermittent SWE benchmark failures and score variabili
github.com
72.0 / 100
as published: 72.0%
INDEPENDENT
counts toward the ranking
glm-5.1
harness Ascend A5 W4A4C8
effort not disclosed
variant verified
vLLM-Ascend project
via Model benchmark dossier
2026-08-15T00:00:00.000Z
facts, not expression
checked: found, but not in a block naming only this model2026-08-30"@context":"https://schema.org","@type":"DiscussionForumPosting","headline":"[DeepSeek-V4-Flash/GLM-5.1] Intermittent SWE benchmark failures and score variabili
github.com
71.4 / 100
as published: 71.4%
INDEPENDENT
counts toward the ranking
glm-5.1
harness Ascend A3 dual-node
effort not disclosed
variant verified
vLLM-Ascend project
via Model benchmark dossier
2026-08-15T00:00:00.000Z
facts, not expression
checked: this figure appears in the page’s row for this model2026-08-30GLM-5.1: A3 dual-node: 71.4%
github.com
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.