A single-response evaluation scores one model output on its own, usually against a predefined set of criteria.
For example, an evaluator might review a customer support response and score whether it is accurate, complete, helpful, and compliant with company policy. The score can be binary, categorical, or numerical.
Single-response evaluation is useful when you care about whether a model meets an absolute quality standard. It is especially helpful for launch criteria, compliance requirements, and tasks where every response needs to satisfy a specific bar.
The main challenge is defining that bar clearly. Strong rubrics and evaluator calibration help make these evaluations more consistent.