A model grading its own homework

Why the finder does not score Shipped

Confidence assigned by the reviewer that raised the finding is thrown away and recomputed by an actor that does not know who raised it.

A reviewer prompted to find gaps will find some. Asked how confident it is, it will report high confidence in the things it just argued for — not from dishonesty, but because it has just spent its whole context constructing the case. The number is a measure of how hard it looked, not of whether the finding is real.

So the score is recomputed, blind

The consultant re-scores every finding 0–100 against an anchored rubric, without knowing which specialist raised it. Attribution is what makes a scorer deferential; removing it is most of what makes the second number worth more than the first.