THE FORGE RANKINGS·2026-08-30

Gemini 3.1 Pro Preview

Google DeepMind
17
RANK of 37 ranked
FERROX INDEX69.9 weighted mean of category positions
GRADEB above the field median
EVIDENCE13 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Not separable from GLM-5.2 and GPT-5.6 Luna at the measured 1.78 point band. The order is printed; the gap is not claimed.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

7of 37 measured
48.4
13.1median 38.963.0

Coding

37of 37 measured
22.4
22.4median 50.070.7

Reasoning

15of 37 measured
71.0
18.1median 62.893.0

Preference

10of 15 measured
69.8
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

ALE-Bench

20/35
1161 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0

APEX Agents

12/28
33.5% best 45.0% · Claude Fable 5
APEX CC-BY-4.0

LMArena Agent

20/21
-0.026 best 0.127 · Claude Opus 5
LMArena CC-BY-4.0

FrontierMath

7/21
36.9% best 47.6% · GPT-5.4 (2026-03-…
Epoch AI CC-BY-4.0

LMArena Text (style-controlled)

4/20
1487 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

DeepSWE

17/17
11.7% best 72.8% · Claude Opus 5
deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent high

Humanity's Last Exam

1/17
46.4% best 46.4% · Gemini 3.1 Pro Pr…
Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0

LMArena WebDev

13/16
1446 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

METR time horizons

1/11
6.4h best 6.4h · Gemini 3.1 Pro Pr…
metr.org CC-BY-4.0

Not measured

19/32
  • SWE-bench Verified (Epoch's own run)
  • FrontierCode
  • Terminal-Bench 4.0
  • Aider Polyglot
  • OSWorld
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
APEX Agents agents 33.5% 26.4 to 40.6 APEX CC-BY-4.0 source
LMArena Agent agents -0.026 5.3 to 10.5 LMArena CC-BY-4.0 source
METR time horizons agents 6.4h 67.7 to 81.2 metr.org CC-BY-4.0 source
Terminal-Bench 2 agents 78.4% 74.9 to 81.9 Terminal-Bench v2 Leaderboard CC-BY-4.0 Forge Code source
ALE-Bench coding 1161 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 source
DeepSWE coding 11.7% 10.2 to 13.2 deepswe.datacurve.ai CC-BY-4.0 mini-swe-agent / high source
LMArena Text (style-controlled) preference 1487 83.3 to 84.3 LMArena CC-BY-4.0 source
LMArena WebDev preference 1446 55.1 to 56.5 LMArena CC-BY-4.0 source
ARC-AGI-2 reasoning 77.1% not published ARC Prize CC-BY-4.0 source
FrontierMath reasoning 36.9% 31.4 to 42.4 Epoch AI CC-BY-4.0 source
GPQA Diamond (Epoch's own run) reasoning 94.1% 90.8 to 97.4 Epoch AI CC-BY-4.0 source
Humanity's Last Exam reasoning 46.4% 42.6 to 50.3 Humanity’s Last Exam (CAIS / Scale AI) CC-BY-4.0 source
OTIS Mock AIME 2024-2025 reasoning 95.6% 89.5 to 100.0 Epoch AI CC-BY-4.0 / high source

Configurations rolled up: 7. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: forge-code, forgecode, gemini-cli, mini-swe-agent, tongagents.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.