Claim · agents

Qwen3.8 Maxon OSWorld

to the range across every scoreable measurement we hold
points between the highest and the lowest
the number the board uses: the lower median, a real row with its own provenance
0 declared configurations, from 0 runners
0 / 2 rows whose cited page still shows this figure, refetched and checked by us

All 0 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.

1 of 2 cited pages no longer show the figure we publish. The page was refetched and split at its own structural boundaries, and no block that names this model carries this number. That does not make the measurement wrong: a leaderboard is a living page and a figure recorded when it was there can be gone by the time you follow the link. It does mean the citation, as printed, no longer supports the number, and this is the only place we know of that says so about its own sources.

Vendor numbers nobody has independently checked

Published by the model's own maker, on a variant of this benchmark that no independent party we hold has run. They are recorded and they are shown. They can never move a rank.

VariantClaimed
unknown 86.1
2.0 19.4

Every measurement

2 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
86.1 / 100
as published: 86.1%
CLAIMED
shown, never scored
qwen3.8-max
harness not_disclosed
effort unknown
variant unknown
Alibaba/Qwen
via Model benchmark dossier
2026-08T00:00:00.000Z
facts, not expression
not checked: the page could not be fetched2026-08-30
datacamp.com
19.4 / 100
as published: 19.4%
CLAIMED
shown, never scored
qwen3.8-max
harness not_disclosed
effort unknown
variant 2.0
Alibaba/Qwen
via Model benchmark dossier
2026-08T00:00:00.000Z
facts, not expression
checked: found, but not in a block naming only this model2026-08-30"key":"osWorld2","name":"OSWorld 2.0","fullName":null,"rawScore":19.4,"displayValue":"19.4%","scoreWidth":19.4,"weightedLabel":"Display only","isWeighted":false
benchlm.ai
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.