“Show the judge's prompt, its samples and where humans disagree with it.”
A model may score the answers, and humans will trust it only if they can see how it decides. The prompt the judge runs on, the cases it scored and the verdicts humans overruled are part of the product. A score with no reasoning behind it is a black box, and humans route around black boxes.
The Context Matrix
Where this principle has been tested. Untested cells are honest gaps, not passing grades — seen evidence for one? Tell us.
| Chat | IDE | CLI | Canvas / Doc | Ambient / Background | |
|---|---|---|---|---|---|
| Coding | |||||
| Knowledge Work | ● | ◐ | |||
| Consumer | |||||
| Enterprise Ops | |||||
| Creative |
● supports · ◐ boundary · ✕ violates · blank untested
Linked Patterns
Boundary Conditions
- knowledge-work × canvas-doc × supervised
No judge, so nothing to open: human grades only.
Evidence
Every metric returns a reason beside its score.
The model's explanation sits beside its score; the reviewer agrees or disagrees.
No judge, so nothing to open: human grades only.
The Agents & Humans Briefing
Agentic experience design, coding agents, MCP, and the signals that matter — weekly, free, in about five minutes.
Free. No spam. Unsubscribe anytime.