← Principles
watching

“Grow from five examples to five hundred in one interface.”

Evaluation becomes a habit when it starts small and stays in one place. Five hand-picked cases run the same way as five hundred, with the same comparison and the same rubric. A separate "advanced mode" for real evaluation tells humans the small version was not real, and they stop at the small version.

The Context Matrix

Where this principle has been tested. Untested cells are honest gaps, not passing grades — seen evidence for one? Tell us.

ChatIDECLICanvas / DocAmbient / Background
Coding
Knowledge Work
Consumer
Enterprise Ops
Creative

● supports · ◐ boundary · ✕ violates · blank untested

Linked Patterns

Evidence

supports

Paste a few cases or generate hundreds; the same tab runs both.

knowledge-work × canvas-doc × supervised · Anthropic Console: the Evaluate tab compares prompt versions side by side

supports

A sentence to Loop becomes rows; the same experiments view handles the result.

knowledge-work × canvas-doc × supervised · Braintrust: pick a baseline and each row shows its score change in red or green