Claim · agents

Gemini 2.5 Flashon Terminal-Bench 2

to the range across every scoreable measurement we hold
points between the highest and the lowest
the number the board uses: the lower median, a real row with its own provenance
0 declared configurations, from 0 runners

All 0 scoreable measurements here come from one runner under one declared configuration, so nothing on this page explains why they differ. That is the finding.

Vendor numbers nobody has independently checked

Published by the model's own maker, on a variant of this benchmark that no independent party we hold has run. They are recorded and they are shown. They can never move a rank.

VariantClaimed
no variant stated 17.1

Every measurement

4 rows. Best first. the board takes the LOWER MEDIAN of the scoreable measurements, which is a real row with its own interval and provenance rather than an interpolated average. a vendor self-report is shown beside the independent numbers and is never one of them; it cannot move a rank.

ScoreProvenanceConfigurationWho ran itWhere to check it
17.1 / 100
as published: 17.1%
UNATTRIBUTED
shown, never scored
gemini-2.5-flash
harness Mini-SWE-Agent
Terminal-Bench
via Epoch AI
2025-11-03
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
16.9 / 100
as published: 16.9%
UNATTRIBUTED
shown, never scored
gemini-2.5-flash
harness Terminus 2
Terminal-Bench
via Epoch AI
2025-10-31
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
16.4 / 100
as published: 16.4%
UNATTRIBUTED
shown, never scored
gemini-2.5-flash
harness OpenHands
Terminal-Bench
via Epoch AI
2025-11-02
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
15.4 / 100
as published: 15.4%
UNATTRIBUTED
shown, never scored
gemini-2.5-flash
harness Gemini CLI
Terminal-Bench
via Epoch AI
2025-11-04
CC-BY-4.0
NO_URL2026-08-30
benchmark leaderboard
the leaderboard for this benchmark, not a link to this row
← The Forge Rankings

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e, as of 2026-08-30. where only one party has measured a model on a benchmark, this page says so rather than presenting a single number as a reconciliation.