AI Lab

How we run AI,
in three small games

We don't bet on one model. We pick the right one for each job, plan for it failing, and keep a person in charge. Everything below is real: our actual routing rules, drafts recorded from real models, and the actual steps that built this page.

Every job has a first-choice model and a plan for when it fails. Pick a job, unplug a model and see what takes over.

1. Pick a job

2. Pull a plug and watch it reroute

Runs in Hermes (always-on agent)

  1. main route

    DeepSeek

    V4.1 Flash

    Running

  2. backup route

    DeepSeek

    V4.1 Flash

    On standby

  3. third route

    DeepSeek

    V4.1 Flash

    On standby

The job runs on DeepSeek. Try pulling the plug.
  • Routine work runs on a fast model built for it.
  • If the main route fails, a second route takes the same job.
  • If both fail, a third route keeps it running.

These are the routing rules we run today, checked 2026-09-24. Names change as better models ship.

The models on our bench

  • DeepSeek V4.1 Flash

    Everyday agent work and code

  • GLM 5.3

    Very long documents, read in one pass

  • MiMo V2.6 Flash

    Reading images and scans

  • Kimi K3

    Hardest problems, only when a person asks

  • Qwen 3.8 Flash

    On standby, not in daily routing

  • Claude Opus / Sonnet

    Leads, reviews and writes client-facing words

And the harnesses they run in

  • Hermes

    Our always-on agent

  • OpenCode

    Code work

  • Claude Code

    Leads, reviews, client writing

Logos are trademarks of their owners, shown to identify the tools we use.