On a laptop.
Qwen 3.8 vs Ornith vs Claude
A 27B model on a laptop matched Opus on hard coding and tied it on animated 3D. Here is exactly where it didn't.
Read the reportWe're an independent R&D lab. We test, bench, build and make AI actually work for everyone.
# toposort: gated climb, live trace probe[minimax] 18/18 SOLVED this task: $0.0007 # expr_interp: the task the same probe cracks alone probe[minimax] 16/29 → ensemble → best 21/29 cheap[deepseek] target: parse-precedence 25/29 ACCEPT cheap[deepseek] target: unary-minus 29/29 ACCEPT · green frontier never called.
A real trace from the published benchmark. Re-run it yourself.
We publish our failures. Nobody else will.
A 27B model on a laptop matched Opus on hard coding and tied it on animated 3D. Here is exactly where it didn't.
Read the report
40.8 to 141.6 tokens per second. The build recipe, the sweeps, and the 18-entry trap catalogue. Follow the steps and you land on the same configuration.
Read the manual
Better code than the frontier models wrote. 20 systems, 196 machine checks, and it's not the model. It's a gate.
Read the reportA single benchmark tells you how a model did on one exam. It cannot tell you where a model stands. We aggregate every public measurement we are licensed to republish, 2,343 of them, across 32 benchmarks from 5 independent sources, and rank on the evidence models actually share. Two models are compared only on the benchmarks they both sat. Vendor self-reports are recorded and never scored.
Top 8 of 37 ranked models, as of 2026-08-30. Hide a whole family of benchmarks, refit the board without it, then ask the hidden family which of two models is better: the ordering had already put that one first 82.4% of the time across 2311 held-out pairs, against 51.3% for an arbitrary ordering of the same models. The pairs it gets wrong are listed on the board. Every number names the benchmark, the evaluator and the licence behind it.
One agent, every agent. Commands 19 AI coding agents from a single interface, with 98 assistants, 178 workflows, and 1,974 curated skills.
The laboratory's primary agent. Runs the gated climb natively, and holds Crucible, the cross-provider judge council everything falls back to.
Premium AI models, smarter pricing. One endpoint routes across 60+ frontier models in four lanes. The verified-routing research runs here in production.
It Just F*cking Works. Local-first infrastructure for AI coding agents: shared memory, smart routing, multi-model cross-audits, and disciplined workflow. If your AI codes, IJFW already runs there.
AI remote control for TradingView Desktop. 109 MCP tools driving symbols, indicators, Pine, snapshots, sweeps, replay, and live chart vision. All local, zero cloud calls.
Benchmarks, build recipes, and the occasional correction when we got something wrong. No fluff.
We do not publish work we cannot defend in review.
We do not release tools we would not use ourselves.
We do not put our name on capability we have not stress-tested.
We publish our failures. Nobody else will.
The Ferrox Labs methodologyLed by Sean Donahoe: originator of the empathy gap construct, the Donahoe Loop, the AI Quality Trident, and the Agent Maturity Ladder. Twenty-six years across Silicon Valley and the Texas tech sector. Ferrox Labs is part of Foundry AI.
If the way we work reads like the way you work, write to us.
jobs@ferroxlabs.com