THE FORGE RANKINGS·2026-08-30

MiniMax-M3

MiniMax
30
RANK of 37 ranked
FERROX INDEX40.6 weighted mean of category positions
GRADED back of the field
EVIDENCE9 / 32 benchmarks measured
QUALITYTHIN Correlated evidence in: preference. Those category scores rest on benchmarks from a single family.

Not separable from o3 (2025-04-16) at the measured 1.78 point band. The order is printed; the gap is not claimed.

Compared to what

Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.

Agents

18of 37 measured
41.3
13.1median 38.963.0

Coding

34of 37 measured
35.9
22.4median 50.070.7

Reasoning

25of 37 measured
54.0
18.1median 62.893.0

Preference

12of 15 measured
69.2
53.1median 71.882.5

Every measurement, and the field behind it

One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.

ALE-Bench

34/35
640 best 2177 · GPT-5.6 Sol
ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0

LMArena Agent

21/21
-0.028 best 0.127 · Claude Opus 5
LMArena CC-BY-4.0

LMArena Text (style-controlled)

15/20
1442 best 1507 · Claude Fable 5
LMArena CC-BY-4.0

FrontierCode

17/17
14.7% best 53.5% · Claude Fable 5
cognition.com CC-BY-4.0 mini-swe-agent none

LMArena WebDev

9/16
1488 best 1626 · Claude Fable 5
LMArena CC-BY-4.0

Not measured

23/32
  • ARC-AGI-2
  • APEX Agents
  • FrontierMath
  • DeepSWE
  • Humanity's Last Exam
  • Terminal-Bench 2
  • METR time horizons
  • Terminal-Bench 4.0
  • Aider Polyglot
  • Cybench
  • BrowseComp
  • IMOAnswerBench
  • SWE-bench Multilingual
  • AIME
  • CyberGym
  • HMMT
  • MCP-Atlas
  • SWE-bench Pro
  • Tau2-Bench Airline
  • Tau2-Bench Banking
  • Tau2-Bench Retail
  • Tau2-Bench Telecom
  • Tool-Decathlon
A missing benchmark is not a zero and is never scored as one.

Check it yourself

Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.

BenchmarkClassPublishedInterval (normalised) Measured byLicenceHarness / effort
LMArena Agent agents -0.028 4.4 to 10.3 LMArena CC-BY-4.0 source
OSWorld agents 75.2% not published OSWorld (XLANG Lab) none stated General model / 100 steps source
ALE-Bench coding 640 not published ALE-Bench (Sakana AI with AtCoder) CC-BY-4.0 source
FrontierCode coding 14.7% not published cognition.com CC-BY-4.0 mini-swe-agent / none source
SWE-bench Verified (Epoch's own run) coding 74.8% not published ambientlight CC-BY-4.0 mini-swe-agent 2.4.2 / enabled source
LMArena Text (style-controlled) preference 1442 76.8 to 78.1 LMArena CC-BY-4.0 source
LMArena WebDev preference 1488 60.1 to 61.8 LMArena CC-BY-4.0 source
GPQA Diamond (Epoch's own run) reasoning 81.3% 75.9 to 86.8 Epoch AI CC-BY-4.0 / none source
OTIS Mock AIME 2024-2025 reasoning 26.7% 13.6 to 39.7 Epoch AI CC-BY-4.0 / none source

Configurations rolled up: 5. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: general-model, mini-swe-agent, mini-swe-agent-2-4.2.

Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.