← Field Notes
MAY 8 · Paper · via arXiv Case DrilldownEval Set Authoring

A paper gives reviewers the number of labels a judge needs

This is the number behind "calibrate with a few labels". A product can tell the reviewer how many cases to label before the judge can be trusted. Not all of them, and not none.

Machine summary of the source

The paper treats the judge as a helper, not a replacement. The judge scores each case, humans score a subsample, and an estimator combines the two so the result is unbiased. The method gives the number of human labels needed for a target level of confidence. Fewer labels are needed when the judge agrees with humans more often.

The summary above is generated; the note at the top is the editorial judgment. Primary source ↗