Each rail is one classification. Every ranked model measured on it is a tick; this model is the orange marker. A hatched rail means this model has no measurement in that classification and its weight was redistributed.
One card per benchmark this model has been measured on. The rail shows where it sits against every other ranked model measured on the same benchmark. Where the evaluator keys on a harness, the harness is named, because the same model scores differently under a different one.
Every number on this page, with who measured it, the interval they published, the harness it was run under, and where to go and read it. Nothing here was measured by Ferrox Labs.
| Benchmark | Class | Published | Interval (normalised) | Measured by | Licence | Harness / effort | |
|---|---|---|---|---|---|---|---|
| BrowseComp | agents | 59.4% | not published | NVIDIA | facts, not expression | Tavily plus terminal workspace / model-card defaults | |
| LMArena Agent | agents | -0.008 | 11.2 to 16.9 | LMArena | CC-BY-4.0 | source | |
| Terminal-Bench 2 | agents | 59.3% | not published | NVIDIA | CC-BY-4.0 | NeMo Evaluator SDK; Harbor / model-card defaults | source |
| ALE-Bench | coding | 887 | not published | ALE-Bench (Sakana AI with AtCoder) | CC-BY-4.0 | source | |
| SWE-bench Multilingual | coding | 74.8% | not published | NVIDIA | facts, not expression | NeMo Evaluator SDK / model-card defaults | |
| SWE-bench Verified (Epoch's own run) | coding | 73.2% | not published | vLLM-Ascend project | CC-BY-4.0 | GPU baseline; 10 groups over 3 rounds / not disclosed | source |
| LMArena Text (style-controlled) | preference | 1466 | 80.3 to 81.5 | LMArena | CC-BY-4.0 | source | |
| LMArena WebDev | preference | 1509 | 62.7 to 64.5 | LMArena | CC-BY-4.0 | source | |
| FrontierMath | reasoning | 33.4% | 28.0 to 38.9 | Epoch AI | CC-BY-4.0 | source | |
| GPQA Diamond (Epoch's own run) | reasoning | 89.9% | 85.7 to 94.1 | Epoch AI | CC-BY-4.0 | source | |
| Humanity's Last Exam | reasoning | 27.2% | not published | NVIDIA | CC-BY-4.0 | NeMo Evaluator SDK / model-card defaults | source |
| IMOAnswerBench | reasoning | 86.8% | not published | NVIDIA | facts, not expression | NeMo Evaluator SDK / model-card defaults | |
| OTIS Mock AIME 2024-2025 | reasoning | 93.3% | 86.0 to 100.0 | Epoch AI | CC-BY-4.0 | source |
Configurations rolled up: 7. Rule: median observed configuration per benchmark (lower median, always a real measurement). Harnesses seen: ascend-a3-dual-node, ascend-a5-w4a4c8, gpu-baseline;-10-groups-over-3-rounds, nemo-evaluator-sdk, nemo-evaluator-sdk;-harbor, tavily-plus-terminal-workspace.
Snapshot 2026-08-30-62e4e85f043b, manifest dc46131315046f1e. Grades are positional: position in the measured field, as a percentile of rank among ranked entries, n=37.