Start here

Overview Shipped

Continually tests and improves the installed teams across LLM backends.

The run observatory: executes every team's eval scenarios as whole runs pinned to one backend per run (claude, codex, qwen), scores reliability with deterministic probes, samples rubric-judged quality, mines the corpus for systemic pipeline defects, and feeds the evidence into teambuilder's improve loop.

Where to go from here

Start with What it is for the one-screen version, Commands for what you can type.