Claim · agents

Jedi 7B W/ GPT-4oon OSWorld

26.8 to 29.3 the range across every scoreable measurement we hold
2.5 points between the highest and the lowest
27.0 the number the board uses: the lower median, a real row with its own provenance
3 declared configurations, from 1 runner

All 3 scoreable measurements here come from ONE runner, so this is not a reconciliation and is not presented as one. What it shows is how far a single runner's number moves when the model is run differently.

What the vendor said, and what someone else measured

Matched on which variant of the benchmark was sat, and on nothing else: the vendor ran it on their harness and the independent party on theirs, and that difference is the comparison rather than an obstacle to it. A vendor number with no independent run of the same variant carries no delta, because a number that cannot be compared honestly must not be compared.

VariantVendor claimsIndependently measuredDifference
as published, no variant stated 27.0 27.0 from 3 runs +0.0

Every measurement

6 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
29.3 / 100
as published: 29.3%
INDEPENDENT
counts toward the ranking
Jedi-7B w/ gpt-4o
harness Agentic framework
effort 100 steps
OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
27.0 / 100
as published: 27.0%
CLAIMED
shown, never scored
Jedi-7B w/ GPT-4o (100 steps) OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
27.0 / 100
as published: 27.0%
INDEPENDENT
counts toward the ranking
Jedi-7B w/ gpt-4o
harness Agentic framework
effort 50 steps
OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
26.8 / 100
as published: 26.8%
INDEPENDENT
counts toward the ranking
Jedi-7B w/ gpt-4o
harness Agentic framework
effort 15 steps
OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
25.0 / 100
as published: 25.0%
CLAIMED
shown, never scored
Jedi-7B w/ GPT-4o (50 steps) OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
22.7 / 100
as published: 22.7%
CLAIMED
shown, never scored
Jedi-7B w/ GPT-4o (15 steps) OSWorld (XLANG Lab)
via OSWorld leaderboard
none stated
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.