NEW: The full Forge catalogue is on the web. Nine new manuals, teardowns, and guides. Browse the library →
Where AI meets the road · Est. 2026

Intelligence is cheap.
Trust is engineered.

We're an independent R&D lab. We test, bench, build and make AI actually work for everyone.

gated-climb.log
# toposort: gated climb, live trace
probe[minimax]   18/18 SOLVED  this task: $0.0007

# expr_interp: the task the same probe cracks alone
probe[minimax]   16/29 → ensemble → best 21/29
cheap[deepseek] target: parse-precedence 25/29 ACCEPT
cheap[deepseek] target: unary-minus 29/29 ACCEPT · green
frontier never called.

A real trace from the published benchmark. Re-run it yourself.

From the lab / nerd fuel Four numbers from the two latest experiments. Failures included.
196 machine checks across 5 tasks 63 hidden from every builder: the overfit detector
133 visible 63 hidden
Read: Gate Engineering →
$0.0063 per accepted task, gated climb Frontier lanes at the same 100%: 2.5x-28.9x more
Anvil · $0.0063
Cheapest frontier · $0.016 · 2.5x
Priciest that finished · $0.1829 · 28.9x
Most expensive model · never finished DNF
Read: Gate Engineering →
141.6 tok/s from a laptop GPU 3.5x the 40.8 stock baseline. Same laptop, config only
Stock · 40.8
Tuned · 141.6
Read: Field Manual No.12 →
256K context window in 24 GB VRAM Measured: 23,504 of 24,462 MB. 958 MB to spare
958 MB
Read: Field Manual No.12 →
§ 01

In the forge

We publish what we find
The Forge / Field Manual No.12 40.8 141.6
tok/s.
Field Manual No.12 · Aug 2026

Maximising Qwen3.8 on a 5090 Laptop

40.8 to 141.6 tokens per second. The build recipe, the sweeps, and the 18-entry trap catalogue. Follow the steps and you land on the same configuration.

Read the manual
Ferrox Labs / Field Notes No.1 100% at 1/29th
the cost.
Field Notes No.1 · Jul 2026

Gate Engineering

Better code than the frontier models wrote. 20 systems, 196 machine checks, and it's not the model. It's a gate.

Read the report
§ 02

The Forge Rankings

One board, every public measurement

A single benchmark tells you how a model did on one exam. It cannot tell you where a model stands. We aggregate every public measurement we are licensed to republish, 2,343 of them, across 32 benchmarks from 5 independent sources, and rank on the evidence models actually share. Two models are compared only on the benchmarks they both sat. Vendor self-reports are recorded and never scored.

Top 8 of 37 ranked models, as of 2026-08-30. Hide a whole family of benchmarks, refit the board without it, then ask the hidden family which of two models is better: the ordering had already put that one first 82.4% of the time across 2311 held-out pairs, against 51.3% for an arbitrary ordering of the same models. The pairs it gets wrong are listed on the board. Every number names the benchmark, the evaluator and the licence behind it.

§ 03

Built here

Research that does not ship is opinion
Wayland mark

Wayland

Desktop agent · Free & open source
The Wayland desktop app

One agent, every agent. Commands 19 AI coding agents from a single interface, with 98 assistants, 178 workflows, and 1,974 curated skills.

Flux Router mark

Flux Router

Model router · Owned & operated
The Flux Router lanes dashboard

Premium AI models, smarter pricing. One endpoint routes across 60+ frontier models in four lanes. The verified-routing research runs here in production.

IJFW

Agent infrastructure · Local-first · Open source
The IJFW wordmark: It Just F*cking Works

It Just F*cking Works. Local-first infrastructure for AI coding agents: shared memory, smart routing, multi-model cross-audits, and disciplined workflow. If your AI codes, IJFW already runs there.

TVControl

TradingView MCP system · All local
The TVControl wordmark: automate TradingView in the language you already speak

AI remote control for TradingView Desktop. 109 MCP tools driving symbols, indicators, Pine, snapshots, sweeps, replay, and live chart vision. All local, zero cloud calls.

Dispatches from the forge

New manuals land in your inbox first.

Benchmarks, build recipes, and the occasional correction when we got something wrong. No fluff.

One list. Unsubscribe any time. We measure our own open rates too.

We do not publish work we cannot defend in review.

We do not release tools we would not use ourselves.

We do not put our name on capability we have not stress-tested.

We publish our failures. Nobody else will.

The Ferrox Labs methodology
§ 04

The laboratory

Forged in the dark

Led by Sean Donahoe: originator of the empathy gap construct, the Donahoe Loop, the AI Quality Trident, and the Agent Maturity Ladder. Twenty-six years across Silicon Valley and the Texas tech sector. Ferrox Labs is part of Foundry AI.

The laboratory is hiring.

If the way we work reads like the way you work, write to us.

jobs@ferroxlabs.com