THE FORGE RANKINGS·2026-08-30

Gemini 2.5 Pro

Google DeepMind
34
RANK of 37 ranked
FERROX INDEX27.7 weighted mean of category positions
GRADEE bottom of the measured field
EVIDENCE10 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

37of 37 measured
13.1
13.1median 38.963.0

Coding

30of 37 measured
40.0
22.4median 50.070.7

Reasoning

27of 37 measured
46.2
18.1median 62.893.0

Preference

15of 15 measured
53.1
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

GPQA Diamond (Epoch's own run)

21/36
85.3% best 94.1% · Gemini 3.1 Pro Pr…
Epoch AI CC-BY-4.0

OTIS Mock AIME 2024-2025

22/36
84.2% best 99.7% · Claude Fable 5
Epoch AI CC-BY-4.0

ALE-Bench

31/35
786 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 32k

APEX Agents

27/28
6.6% best 45.0% · Claude Fable 5
APEX CC-BY-4.0

FrontierMath

16/21
14.1% best 47.6% · GPT-5.4 (2026-03-…
Epoch AI CC-BY-4.0

SWE-bench Verified (Epoch's own run)

21/21
57.6% best 83.5% · Claude Opus 4.7
Epoch AI CC-BY-4.0

LMArena Text (style-controlled)

14/20
1446 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

LMArena WebDev

16/16
1226 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

Not measured

22/32
  • LMArena Agent
  • DeepSWE
  • FrontierCode
  • Humanity's Last Exam
  • METR time horizons
  • Terminal-Bench 4.0
  • Aider Polyglot
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 6.6% 3.7 to 9.5 APEX CC-BY-4.0 source
Terminal-Bench 2 agents 19.6% 13.9 to 25.3 Terminal-Bench v2 Leaderboard CC-BY-4.0 Gemini CLI source
ALE-Bench coding 786 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 / 32k source
SWE-bench Verified (Epoch's own run) coding 57.6% 53.1 to 62.0 Epoch AI CC-BY-4.0 source
LMArena Text (style-controlled) preference 1446 77.6 to 78.3 LMArena CC-BY-4.0 source
LMArena WebDev preference 1226 26.2 to 30.2 LMArena CC-BY-4.0 source
ARC-AGI-2 reasoning 4.0% not published ARC Prize CC-BY-4.0 / 16k source
FrontierMath reasoning 14.1% 10.1 to 18.1 Epoch AI CC-BY-4.0 source
GPQA Diamond (Epoch's own run) reasoning 85.3% 81.1 to 89.4 Epoch AI CC-BY-4.0 source
OTIS Mock AIME 2024-2025 reasoning 84.2% 74.7 to 93.6 Epoch AI CC-BY-4.0 source

Configurations rolled up: 8. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: gemini-cli, mini-swe-agent, openhands, terminus-2.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.