Overview Shipped
Continually tests and improves the installed teams across LLM backends.
The run observatory: executes every team's eval scenarios as whole runs pinned to one backend per run (claude, codex, qwen), scores reliability with deterministic probes, samples rubric-judged quality, mines the corpus for systemic pipeline defects, and feeds the evidence into teambuilder's improve loop.
Where to go from here
Start with What it is for the one-screen version, Commands for what you can type.