← Principles
watching

“Show the judge's prompt, its samples and where humans disagree with it.”

A model may score the answers, and humans will trust it only if they can see how it decides. The prompt the judge runs on, the cases it scored and the verdicts humans overruled are part of the product. A score with no reasoning behind it is a black box, and humans route around black boxes.

The Context Matrix

Where this principle has been tested. Untested cells are honest gaps, not passing grades — seen evidence for one? Tell us.

ChatIDECLICanvas / DocAmbient / Background
Coding
Knowledge Work
Consumer
Enterprise Ops
Creative

● supports · ◐ boundary · ✕ violates · blank untested

Linked Patterns

Boundary Conditions

  • knowledge-work × canvas-doc × supervised

    No judge, so nothing to open: human grades only.

Evidence

supports

Every metric returns a reason beside its score.

knowledge-work × IDE × supervised · DeepEval: each metric explains itself, and regressed rows turn red

supports

The model's explanation sits beside its score; the reviewer agrees or disagrees.

knowledge-work × canvas-doc × supervised · Freeplay: agree or disagree with the judge, ten examples at a time

boundary

No judge, so nothing to open: human grades only.

knowledge-work × canvas-doc × supervised · Anthropic Console: the Evaluate tab compares prompt versions side by side