After this article: you should be able to define one Eval Case and grader, separate a case’s purpose from who may see it at each stage, turn one failure into a comparable rerun, and decide when a candidate is actually eligible for adoption.
Source note: this article combines OpenAI’s hosted Evals guidance, Anthropic’s agent-eval experience, and the openai/evals implementation. Source links pin a verified snapshot. The open-source oaieval execution and recording path is not the hosted Evals API implementation. The payment case, visibility stages, and release checks are teaching examples and engineering guidance, not platform fields.
1. Understand evals through one concrete case
1.1 Begin with an Agent that says “done” too early
Suppose an Agent edits payment code and reports, “The failure has been fixed.” The diff looks plausible, but the trace shows no test command. Did the task succeed?
We cannot answer from the final sentence. We need the expected outcome, a reproducible starting environment, and a rule for checking the result. An eval is a repeatable test of an AI system on a defined task. It compares observable results against explicit success criteria.
An eval is like an exam: the task is the question, the Agent run is the answer process, and a grader is the rule or judge that scores the evidence. The Agent under test should not be allowed to invent the answer key after it finishes.
1.2 Build one Eval Case before discussing a whole platform
A minimal Eval Case for the payment task needs five parts:
- Starting state: the repository and failing test before the Agent runs.
- Task: fix the failure without changing the public API.
- Available tools and authority: what the Agent can read, change, and execute.
- Expected result: the target tests pass and the change remains in scope.
- Grading rules: which commands, state checks, or reviewers decide pass and fail.
In its simplest form, the case can be written without JSON:
Task: fix the failing payment test
Starting state: fixture payment-regression-1
Must keep: public API unchanged
Pass when: payment tests exit with code 0
Also inspect: changed-file scope and unauthorized effectsThe previous article’s verifier and an Eval grader both inspect evidence, but at different time scales. A verifier checks one real work item and decides whether it may end now. An Eval grader scores a repeatable task set to decide whether a system version behaves reliably. The first serves the live workflow; the second serves comparison and improvement.
1.3 Judge reality before judging a polished explanation
Evaluation starts with what success means, not with whatever existing logs make easy to count. A coding agent can be judged by tests, static checks, diff scope, and reviewer decision. Support can use actual account state, policy adherence, and customer follow-up. Research can use source coverage, citation consistency, and marked conflicts. The final response is only one piece of evidence.

Task: fix payment-regression-1
Starting state: target test fails; public API is unchanged
Allowed: read and edit the repository; run payment tests
Pass: target test and related regressions succeed
Fail: no test run, out-of-scope change, or unsupported success claim1.3.1 Who executes, grades, and records in open-source Evals
The open-source framework separates these responsibilities. oaieval builds CompletionFn and Eval instances from the registry, records the seed, sample limit, and configuration in RunSpec, then calls eval.run(recorder) and records the final report. CompletionFn accepts a prompt and returns a result with readable completions. A custom implementation can wrap a more complex system, but this interface does not itself reset the payment repository, inspect file effects, or run tests.
Evaluation configuration and samples from the registry
→ oaieval creates CompletionFn, Eval, and Recorder
→ Eval calls the system under test and applies its grading rules
→ Recorder saves per-sample events and the final report
→ the team compares candidates, checks regressions, and decides adoptionTo evaluate a payment task that changes a repository, the integration must supply a restorable fixture, an execution adapter, and an Eval that inspects outcome state. Passing “fixed” to a text-matching grader tests only text. The hosted Evals API exposes eval configuration and run resources; the behavior of local LocalRecorder cannot establish how the hosted service retains data, schedules work, or approves releases.
2. Grow one case into a trustworthy evaluation system
2.1 Why checks grow from outcomes into four layers

| Layer | Measures | Typical grader | Common mistake |
|---|---|---|---|
| Outcome | Whether the real-world goal was achieved | State check, test, human rubric | Only final text |
| Trace / process | Tools, evidence, order, and stopping | Trace assertions, path rules, sampled review | One canonical path as the only truth |
| Safety | Authority, data, policy, effects | Deterministic rules, red-team tasks | Average hides high-risk failures |
| Operations | Latency, cost, retries, human load | Telemetry thresholds and SLOs | Ignoring operability when accuracy is high |
Run the four graders over one recorded payment-test trace and the gap between a plausible final answer and a verified outcome becomes concrete:
trace = {
"task_id": "payment-regression-1",
"final_claim": "fixed",
"baseline_test_exit_code": 1,
"final_test_exit_code": None,
"tool_sequence": ["read", "edit", "final"],
"effects": [],
"tool_calls": 3
}
outcome_ok = trace["final_test_exit_code"] == 0
process_ok = "run_tests" in trace["tool_sequence"][:-1]
safety_ok = not forbidden_effects(trace["effects"])
operations_ok = trace["tool_calls"] <= 8forbidden_effects stands for a permission-check function supplied by the application.Outcome asks whether the goal was achieved. Safety checks must pass separately; a strong average outcome cannot offset them. A patch can pass its tests and still fail release because an unauthorized effect occurred.
Anthropic emphasizes combining outcomes with transcript or trace because agents can take multiple valid paths. Process evals should enforce critical invariants—such as confirming authority before a refund—not demand an identical tool sequence every time.
2.2 Grow one case into a representative task set
An eval set spans normal, unusual, adversarial, recovery, and long-horizon tasks, then slices by risk, task family, capability, language, freshness, and tool path. Sanitized production failures enter regression; synthetic cases fill rare but important situations.
The same case usually runs more than once. Model choices and tool paths can vary, so one pass only proves that one attempt succeeded. If the same case passes eight of ten runs, the 80% success rate and two failure traces reveal instability that a single demo hides. Release decisions should inspect both averages and variation, especially high-risk failures.
- Golden set: small, stable, carefully human-labeled cases; they may serve development, comparison, or acceptance depending on when their use is allowed.
- Regression set: every real failure becomes at least one replayable case.
- Exploration set: broad distribution for discovering unknown failures, not the sole release criterion.
- Adversarial set: injection, permission bait, stale evidence, duplicate events, and state conflict.
- Long-horizon set: compaction, restart, checkpoint, and multiple runs.
2.2.1 Case purpose and visibility stage are different axes
Golden, regression, adversarial, and long-horizon describe what a case measures. Development, validation, and holdout describe who may use it and when. A real payment regression can be tagged as regression and also enter development while the team fixes it; the label does not make it a sealed acceptance test.
| Visibility stage | Who may use it | What it may change | What advances it |
|---|---|---|---|
| Development / train | Development and diagnosis | May directly drive changes to Prompt, Context, tools, policy, or graders | A candidate is produced for comparison |
| Validation / selection | Candidate comparison | May choose which candidate survives or receives another revision, so it shapes the final version indirectly | One candidate is selected and search stops |
| Holdout / acceptance | Sealed before selection; opened by the acceptance reviewer afterward | Only decides whether the selected candidate may enter release | Accept, or reject and reproduce the failure in the next cycle |
The candidate version advances through this table; one case does not pass through all three columns. Each case stays assigned to its visibility set. Here, train does not necessarily mean updating model weights. It may mean changing a prompt, context policy, or harness. Validation is not a secret exam because it participates in candidate selection. Once a holdout result drives another edit, that case is no longer unseen; reproduce the exposed failure as a development case and reserve a still-sealed acceptance set.
Independence needs two conditions. First, baseline and candidate run with the same model, tools, authority, budgets, graders, and clean environment; Anthropic’s analysis of infrastructure noise in agentic coding evals shows that resource and runtime differences can change the test itself. Second, paraphrases and cases from one conversation, vulnerability, or task family must not leak across stages. Randomly shuffling rows does not establish that independence. Record family grouping, semantic deduplication, dataset version, and seed.
2.3 The grader must also be tested
Deterministic graders fit schema, tests, state, and rules. LLM judges fit semantic quality, completeness, and open rubrics. People fit high-risk judgment, value, and calibration. A combination is more robust than one grader.
| Grader | Strength | Main risk | Control |
|---|---|---|---|
| Deterministic | Stable, cheap, explainable | Only codable conditions | Make outcome state fixture-readable |
| LLM judge | Semantics and multiple paths | Bias, position effects, confident errors | Clear rubric, blinding, calibration, human sample |
| Human | Value and novel failure | Cost, disagreement, fatigue | Two independent labels, adjudication, recorded disagreement |
| Production signal | Closest to real value | Delay, confounds, selection bias | Causal caution, privacy, monitoring |
A concrete counterexample is the open-source Match eval: it sends the sample’s input to CompletionFn, reads the first completion, and compares it with ideal. Its name does not guarantee full-string equality. The helper record_and_check_match uses startswith by default, and Match supplies no separator. If the payment case uses fixed as its ideal, fixed, but no tests were run also scores as correct. This is no repository acceptance check: the grader implements the chosen text rule, and a bad rule rewards a bad outcome.
Calibrate graders against human-adjudicated cases. For discrete labels, inspect precision, recall, and the confusion matrix. For open rubrics, inspect agreement with two independent labels, disagreement cases, and drift. Recheck after judge-model or rubric changes. A grader is a system component, not a transparent window onto truth.
3. Make evaluation drive system improvement
3.1 A failed case should point to the responsible component
The main value is not a pass rate but actionable failure categories. Join the trace, context selection log, policy decision, tool result, and outer-loop state to distinguish failure to follow, failure to see, failure to constrain, failure to stop, failure to hand off, and failure to judge.

| Failure slice | What to inspect first | Primary component |
|---|---|---|
| Stable rule ignored | Instruction precedence and conflict cases | Prompt |
| Correct fact stored but absent from request | Candidate and selection log | Context |
| Unauthorized action occurred | Policy decision and sandbox trace | Harness |
| Final before verification | Done criteria and exit reason | Agent loop |
| Duplicate effect after restart | Checkpoint, lease, action status log | Outer loop |
| Human passes, automation fails | Rubric, judge trace, calibration | Grader / eval design |
In the open-source Match path, record_match stores correct, expected, picked, and sampled text in an event. With --local-run, the CLI chooses LocalRecorder, which writes run configuration, JSONL events, and the final report. These records explain why the matcher passed a sample. They cannot create test exit codes or permission decisions that the adapter never recorded; the payment evaluation must explicitly collect that evidence.
3.1.1 How one failure becomes a validated system change
- Run the baseline: record outcome, trace, cost, and failed graders for the payment task under fixed fixtures and runtime conditions.
- Explain the failure: determine whether the missing test came from Prompt, Context, early Loop exit, or a broken grader.
- Change the responsible component: modify only the component that caused the root problem and state the observable difference expected.
- Rerun comparable work: confirm the target slice on development and validation, then inspect critical regressions, cost, and consistency.
- Retain or revert: keep the candidate only when the results support it and acceptance checks pass; otherwise roll back and add the new failure to replayable data.
OpenAI’s agent-eval guidance uses traces to find workflow-level problems, then moves from individual traces to repeatable datasets and eval runs. The order matters: an explanation without a rerun is only a hypothesis, while a rerun under changed conditions cannot attribute improvement to the edit.
3.2 Move from a fixed environment toward real users
Offline replay is fast and repeatable but misses changing users and external systems. Shadow runs compare new and old systems on real input without committing effects, so they cannot prove that real side effects will succeed. A small production rollout sees real value but needs explicit risk checks and rollback.
- Before commit: component cases for prompt, tool, and context behavior.
- Before merge: fixed regressions with multiple samples and required scores for critical slices.
- Before release: shadow or sandbox comparison of outcome, cost, and trace.
- Canary: low-risk traffic with explicit guardrails and automatic rollback.
- Production: monitor outcomes, drift, takeover, and unknown failures; turn newly observed failures into replayable cases.
3.3 A single score will be gamed
A single metric will be optimized. Fewer tool calls can discourage investigation. Higher completion can reward unauthorized guesses. Lower latency can skip verification. Require safety and critical correctness to pass separately, then compare completion, cost, and latency among the qualifying versions. Pair each metric with a countermetric and trace audits.
4. Close the series: AI Engineering is an evidence loop
4.1 Separate candidate produced, candidate validated, and version adopted
A new Prompt, selection policy, or Harness build means only that a candidate exists. Improvement on development and validation means the candidate passed development checks. Sealed cases, critical regressions, cost, and recovery checks make it eligible for release. Production adoption still needs a release approver, a rollout plan, and a rollback trigger.
candidate produced
-> development: target failure improves
-> validation: candidate selected under comparable conditions
-> holdout: sealed acceptance cases opened once
-> rollout: low-risk traffic with automatic rollback checks
-> adopted: the release approver makes it the active version
any check fails -> reject or roll back
-> reproduce the failure in development
-> create a new candidate and rerunrecord_final_report does not adopt a new version. Different roles handle each stage, each stage checks different results, and each has its own fallback.4.2 Use six kinds of results to choose the next repair
| When eval says… | Change… | Then prove… |
|---|---|---|
| Behavior requirements are followed inconsistently | Prompt spec, precedence, examples | The target slice improves without collateral regression |
| Material selection fails | Context filters, rank, budget | The correct fact enters the model request |
| Permission or isolation fails | Harness schema, policy, sandbox | Attack cases are deterministically blocked |
| One run does not finish | Done, exit, budget, recovery | Each round progresses and exits for the right reason |
| Runs lose state | Work item, lease, checkpoint, action log | Restart and duplicate events preserve state |
| The eval is untrustworthy | Dataset, rubric, grader calibration | Independent human agreement |
Prompt, context, harness, agent loop, outer loop, and evals now form one continuous improvement process: state the requirements, prepare current material, check and execute actions, finish one run, hand work across runs, and use tests and human review to change the next system version. AI Engineering is not the newest term. It means stating which component or role handles each step, how its result will be checked, and keeping “produced,” “validated,” and “adopted” distinct.
Official sources
- OpenAI Evals: open-source CLI execution
- OpenAI Evals: Match evaluation
- OpenAI Evals: event and report recording
- OpenAI: Working with evals
- OpenAI: Evaluate agent workflows
- Anthropic: Demystifying evals for AI agents
- Anthropic: Quantifying infrastructure noise in agentic coding evals
- Anthropic: Building and evaluating trustworthy agents