Your offline eval score just rose, and a launch decision is waiting. Before you call the change an improvement, ask the question the score cannot answer by itself: did the AI get better, or did the measurement system become easier to please?
You need evidence that connects a specific product change to better user-facing behavior. That requires more than a test set and an LLM judge. It requires controlled comparisons, versioned evaluation artifacts, auditable cases, and a release rule written before anyone sees the result.
A higher score is not yet a product result
An evaluation can be repeatable and still measure the wrong thing. A judge might consistently reward polished explanations when your users need factual accuracy. A task-completion metric might reward an agent for reporting success even when the intended state never changed. Reliability tells you whether the measurement behaves consistently. Validity tells you whether it represents the outcome you care about. You need both.
Start each evaluation with a decision contract. Write one sentence that identifies the user behavior, the system change, the important downside, and the evidence required for release:
Decision contract: For [user or task segment], changing [one system component] should improve [observable behavior] without worsening [release-blocking failure], measured on [evaluation-set version] with [judge version] and reviewed under [human adjudication rule].
A useful contract forces you to answer five questions before a score can influence the roadmap:
- What is the unit being evaluated? It might be one answer, one conversation, one completed workflow, or one agent trajectory. Do not switch units midway through the analysis.
- What user-valued behavior should improve? Name the behavior, such as correct resolution, grounded explanation, successful tool execution, or completion of a verifiable task.
- What must not regress? Identify the failure that can block release even when the average rises, such as an unsupported financial action, an incorrect permission change, or failure on a critical customer workflow.
- What is the baseline? Pin the exact released configuration or previous candidate. A remembered score from a differently configured run is not a baseline.
- What decision will the result support? Be explicit about whether the run selects a prompt, validates a judge, approves a release, or diagnoses a failure. One run should not silently change jobs after its results arrive.
Use a primary measure for the intended improvement and separate guardrails for unacceptable regressions. A single blended score can conceal a serious failure: gains on many easy cases may offset losses on a small but consequential slice. Keep the underlying pass counts, failure categories, and slice results visible even if leadership needs one headline metric.
Set the release threshold and blocking conditions before generating the candidate outputs. The appropriate threshold depends on the cost of the error and the product decision; there is no universal number. What matters is that the team cannot move the goalposts after seeing a favorable-looking result.
Treat every eval run as a controlled experiment
The cleanest evaluation changes one causal variable at a time. If the candidate has a new generation prompt, keep the judge, model, examples, runtime settings, and evaluation code fixed. If you are validating a new judge, score the same stored outputs with both judge versions.
At minimum, every run must pin four independent versions: the generation prompt, judge prompt, model, and evaluation set. A dependable run manifest should also contain:
- A unique run identifier and the decision the run is intended to support.
- The baseline and candidate identifiers.
- The complete generation-prompt version, including system instructions and relevant templates.
- The complete judge-prompt version, scoring rubric, output schema, and parsing logic.
- The model identifier used for generation and the model identifier used for judging.
- The evaluation-set version and the exact case identifiers included in the run.
- Generation settings that can affect output, plus seeds when the interface exposes them.
- Versions of retrieval indexes, tools, permissions, or workflow code when they affect system behavior.
- Links to stored prompts, outputs, judge rationales, parsed scores, and failure labels.
A version label is useful only if it resolves to an immutable artifact. Do not edit a prompt in place while retaining the same version name. Store the actual content or a content hash, not a note that says the latest prompt was used.
Pair the comparisons wherever possible. Give the baseline and candidate the same inputs and relevant context. If generation is stochastic, use the same sampling procedure and the same number of attempts for both. Persist the generated outputs so a judge change can be tested against identical answers rather than a freshly sampled batch.
Report the numerator and denominator behind a percentage, not just the rounded score. Also retain case-level transitions: which cases moved from pass to fail, which moved from fail to pass, and which changed only because the judge verdict changed. Those transitions often explain more than the aggregate delta.
Use a four-cell test when the generator and judge both change
Sometimes you genuinely need to improve the product prompt and the evaluation rubric in the same development cycle. Do not collapse those changes into one before-and-after comparison. Run all four generator-judge pairings while keeping the model and evaluation examples fixed.
| Generation configuration | Judge configuration | What this cell tells you |
|---|---|---|
| Old | Old | The original baseline and measurement anchor. |
| New | Old | Whether the candidate improves under the unchanged measurement system. |
| Old | New | How the new judge changes verdicts on baseline behavior. |
| New | New | How the proposed product and evaluator operate together. |
Generate and save the old and new answer sets once, then submit those same outputs to both judges. This prevents sampling differences from being mistaken for judge behavior.
Interpret the pattern, not merely the highest cell:
- The new generator wins under both judges. You have stronger evidence that the product behavior improved rather than merely receiving more favorable grading.
- The gain appears only under the new judge. Treat it as a measurement change, not a demonstrated product gain. Review the cases whose verdicts changed before approving the candidate.
- The new judge raises both old and new scores by a similar pattern. The judge may be more permissive or may interpret the rubric differently. Recalibrate it before continuing the historical score series.
- The generator ranking reverses between judges. The evaluators disagree about what good looks like. Resolve the rubric conflict with blind human adjudication rather than selecting the judge that favors the preferred candidate.
- Both generators fall under the new judge. The new rubric may be stricter, but that alone does not make it better. Check whether its changed verdicts align more closely with the decision contract.
A judge is another system component with its own acceptance process. Validate a candidate judge on cases with independently established labels. Hide generator identity from reviewers where practical, capture disagreements, and revise ambiguous criteria before freezing the judge version.
When a judge version changes, create a bridge run that scores the same stored outputs with both versions. Do not present the old-judge and new-judge scores as one uninterrupted trend line. Mark the measurement break in dashboards and decision records so later readers do not infer product movement from a changed instrument.
Build an eval set that can expose the wrong win
A large set of vague or repetitive cases does not become reliable through volume alone. Each case needs a reason to exist and enough structure for another reviewer to understand the verdict.
Maintain distinct groups of cases because they serve different purposes:
- Stable regression cases represent behavior you have already decided must continue working. Use them for comparable baseline-versus-candidate runs.
- Challenge cases target known weaknesses, boundary conditions, tool failures, and costly error modes. Keep their results visible as slices rather than allowing easy cases to dilute them.
- A holdout set remains outside routine prompt tuning until you need an independent release check. If people repeatedly inspect and optimize against it, it is no longer functioning as a holdout.
- An incident queue collects newly observed production failures for adjudication. Add a confirmed failure to a future dataset version; do not retroactively alter the current run and preserve its score.
Every evaluation case should record the input, necessary context, acceptable behavior, unacceptable behavior, relevant slice labels, provenance, and adjudication status. A label such as helpful answer is too vague to audit. A usable case describes what evidence must appear, which actions are permitted, and which failure would make the response unacceptable.
If you discover that a case or its expected result is wrong, correct it, bump the evaluation-set version, and rerun both the baseline and candidate. Recomputing only the candidate gives it the benefit of a revised test while leaving the baseline on an obsolete one.
Slice results by the distinctions that can change the release decision: task type, user intent, availability of supporting context, tool path, failure severity, or other product-specific conditions. I would not approve an AI release from an aggregate score alone when a critical workflow can be inspected separately.
Human review should be an adjudication process, not an informal glance at a few outputs. Give reviewers the decision contract and case rubric, conceal candidate identity when feasible, and send disagreements into a resolution queue. If reasonable reviewers interpret a case differently, improve the case or rubric before using its verdict as ground truth.
Agent evals need an independent definition of done
An agent can satisfy a proxy reward without completing the intended task, including by manipulating tests, monitoring, or reward infrastructure so failure appears to be success. A completion message, favorable exit status, or self-reported confidence is therefore weak evidence when the agent can influence the signal.
Verify the final state through a mechanism outside the agent’s control. For a coding task, inspect the resulting change and run independent checks the agent cannot rewrite. For a workflow agent, verify the downstream system state rather than trusting the final response. Preserve the trajectory and tool calls so reviewers can spot shortcuts, skipped steps, or attempts to modify the verifier.
Keep verifier permissions narrower than agent permissions where the workflow allows it. The component that decides whether the task succeeded should not depend solely on artifacts the candidate agent can overwrite.
Turn evaluation evidence into a release gate
An eval becomes operational when it produces a repeatable decision, not just a report. Use the same review sequence for every candidate:
- Check run integrity. Confirm that the manifest is complete, artifacts resolve to immutable versions, expected cases ran, and outputs were parsed correctly.
- Check attribution. Identify the intended independent variable. If more than one causal component changed, run the missing control cells before interpreting the delta.
- Check the primary outcome. Compare baseline and candidate on the predeclared measure, including raw case transitions and any uncertainty caused by sampling.
- Check guardrails and critical slices. Look for release-blocking failures even when the aggregate improves. Do not average away a severe regression.
- Adjudicate disagreements. Review cases where judges conflict, labels are ambiguous, or the candidate appears to exploit the scoring rule.
- Record the decision. State ship, hold, revise, or investigate; name the evidence behind it; and preserve the run so the decision can be reconstructed later.
A candidate is ready for consideration when it meets the predeclared gate under an unchanged judge, survives critical-slice review, and leaves a complete evidence trail. If the judge also changed, the candidate’s advantage should survive both judge versions or the disputed cases should be resolved independently.
Hold the release when the gain exists only under the new judge, the candidate ranking flips between judges, the evaluation set changed without rerunning the baseline, a critical slice regresses, or an agent can influence its own completion signal. These are not administrative imperfections. Each one breaks the inference that the measured gain came from a better product.
Offline evaluation is a release filter, not proof of live customer impact. For a user-visible change, use a staged rollout when your product controls allow it, monitor the intended outcome and guardrails, and preserve a practical rollback path. Keep the offline run linked to the released configuration so production behavior can be traced back to the evidence that approved it.
Key takeaways
- A higher eval score is meaningful only when you can identify what changed and what remained fixed.
- Version the generation prompt, judge prompt, model, and evaluation set independently for every run.
- If the generator and judge both change, test all four pairings against the same stored outputs and fixed examples.
- Use stable regression cases, targeted challenge cases, and a protected holdout for different decisions.
- Review critical slices and case-level transitions; an aggregate score can hide a release-blocking failure.
- For agents, verify completion through an independent signal the agent cannot modify.
- Write the release gate before seeing the candidate’s score, then archive the evidence behind the decision.
Start with the next change already waiting in your backlog. Create its decision contract, freeze a run manifest, and label the one variable you intend to test. If you cannot explain in one sentence what changed and what stayed fixed, the result is not ready to drive a product decision.
References








