Claude, Codex, Qwen

Multi-model testing Preview

Backends are pinned per run, so portability defects — prompts that only work on one model — surface as a finding type rather than a mystery.

Three backends are wired: Claude, Codex, and Qwen running locally against Ollama. A campaign cell pins its backend in the invocation itself. Comparative reads are queries rather than prose: score means grouped by team, model, effort and criterion, so the comparison is per criterion rather than per run — the excellence criteria live once in the team's shared intent contract, not per scenario, which is what makes them the axis two models can be lined up along. Claude is the only default — Codex is registered and usable but off while it is hardened.

Variance is a first-class metric

A team scoring 85 with a spread of 3 is better than one scoring 88 with a spread of 20, and the harness is built to say so. Re-judging for drift is append-only — a second judgment never overwrites the first, because the disagreement between them is itself the data.