Continually tests and improves the installed teams across LLM backends.

What it is Shipped

The one-screen version of teamtester: what it is for, what it produces, and what it does not do.

The run observatory: executes every team's eval scenarios as whole runs pinned to one backend per run (claude, codex, qwen), scores reliability with deterministic probes, samples rubric-judged quality, mines the corpus for systemic pipeline defects, and feeds the evidence into teambuilder's improve loop.

At a glance

  • 0 specialists, 0 lenses
  • 6 top-level commands
  • No scopes — the command decides the roster, not a scope flag
  • Produces runs, so it answers the six shared run commands