← Field Notes
JUL 29 · Clipped · via LangChain blog Case DrilldownEval Set Authoring

LangSmith scores its judge against the team's own grades

The alignment score is the trust signal. A judge version with a number beside it, measured against the team's own grades, is something a product manager can decide to rely on.

Machine summary of the source

Align Evals is a workspace for the judge prompt. It shows an alignment score against examples humans graded. A side-by-side of human and model scores sorts to the disagreements. The previous judge version's score stays as a baseline. Corrections are kept as examples the judge learns from. LangChain's own line on the problem: the scores did not match what a person on the team would say.

The summary above is generated; the note at the top is the editorial judgment. Primary source ↗