“Grow from five examples to five hundred in one interface.”
Evaluation becomes a habit when it starts small and stays in one place. Five hand-picked cases run the same way as five hundred, with the same comparison and the same rubric. A separate "advanced mode" for real evaluation tells humans the small version was not real, and they stop at the small version.
The Context Matrix
Where this principle has been tested. Untested cells are honest gaps, not passing grades — seen evidence for one? Tell us.
| Chat | IDE | CLI | Canvas / Doc | Ambient / Background | |
|---|---|---|---|---|---|
| Coding | |||||
| Knowledge Work | ● | ||||
| Consumer | |||||
| Enterprise Ops | |||||
| Creative |
● supports · ◐ boundary · ✕ violates · blank untested
Linked Patterns
Evidence
Paste a few cases or generate hundreds; the same tab runs both.
A sentence to Loop becomes rows; the same experiments view handles the result.
The Agents & Humans Briefing
Agentic experience design, coding agents, MCP, and the signals that matter — weekly, free, in about five minutes.
Free. No spam. Unsubscribe anytime.