AI Lab
How we run AI,
in three small games
We don't bet on one model. We pick the right one for each job, plan for it failing, and keep a person in charge. Everything below is real: our actual routing rules, drafts recorded from real models, and the actual steps that built this page.
Every job has a first-choice model and a plan for when it fails. Pick a job, unplug a model and see what takes over.
1. Pick a job
2. Pull a plug and watch it reroute
Runs in Hermes (always-on agent)
- main route
DeepSeek
V4.1 Flash
Running
- backup route
DeepSeek
V4.1 Flash
On standby
- third route
DeepSeek
V4.1 Flash
On standby
- Routine work runs on a fast model built for it.
- If the main route fails, a second route takes the same job.
- If both fail, a third route keeps it running.
These are the routing rules we run today, checked 2026-09-24. Names change as better models ship.
The models on our bench
DeepSeek V4.1 Flash
Everyday agent work and code
GLM 5.3
Very long documents, read in one pass
MiMo V2.6 Flash
Reading images and scans
Kimi K3
Hardest problems, only when a person asks
Qwen 3.8 Flash
On standby, not in daily routing
Claude Opus / Sonnet
Leads, reviews and writes client-facing words
And the harnesses they run in
Hermes
Our always-on agent
OpenCode
Code work
Claude Code
Leads, reviews, client writing
Logos are trademarks of their owners, shown to identify the tools we use.