← Principles
watching

“Calibrate the judge with a few human labels and show the disagreement.”

Labeling each case is a research project; labeling none leaves the judge unchecked. A handful of human labels on the cases the judge is least sure about tunes it. The disagreement rate then tells the team how far to trust the rest. Disagreement kept in view is the trust signal.

The Context Matrix

Where this principle has been tested. Untested cells are honest gaps, not passing grades — seen evidence for one? Tell us.

ChatIDECLICanvas / DocAmbient / Background
Coding
Knowledge Work
Consumer
Enterprise Ops
Creative

● supports · ◐ boundary · ✕ violates · blank untested

Linked Patterns

Evidence

supports

An alignment score against human-graded examples, per judge version.

knowledge-work × canvas-doc × supervised · LangSmith Align Evals: tune the judge against human grades and keep the score

supports

Ten examples at a time; an alignment score decides when the judge is trusted.

knowledge-work × canvas-doc × supervised · Freeplay: agree or disagree with the judge, ten examples at a time