How-to guide
DraftDesign an evaluation the agent cannot edit
Separate evaluation assets and authority from the agent so a proposal cannot redefine the test that decides whether it passes.
- For
- Teams building or operating agent workflows
- You will finish with
- A protected evaluation with explicit ownership, fixed inputs, reproducible scoring, and a tested failure path.
Before you start
- A bounded task the agent already performs or will perform
- A small set of representative inputs and expected outcomes
- Control over the agent's tools, credentials, and execution environment
An evaluation is only independent when the agent cannot change the inputs, scoring rules, or pass decision during the run. A prompt that says “do not edit the tests” is an instruction. It is not a boundary.
This guide builds that boundary around one workflow. Start small. Protect a single decision before trying to design an evaluation platform.
1. State the decision
Write the question the evaluation must answer. Keep it narrower than “is the agent good?”
For a coding agent, the decision might be:
May proposal
Preplace baselineBwithout breaking any protected behavior?
Define the possible verdicts before writing the scorer:
PASS all required checks ran and the proposal met the acceptance rules
REGRESSION one or more protected cases became worse than the baseline
INVALID the submission or evidence was malformed
NO_VERDICT the evaluator could not complete the required checks
Do not turn NO_VERDICT into PASS. A missing decision should stop promotion.
2. Draw the write boundary
List every asset that can change the verdict. Record who can read it, who can write it, and what identity performs the write.
| Asset | Agent access | Evaluation owner | Required control |
|---|---|---|---|
| Public rubric | Read | Evaluation maintainers | Agent has no write path |
| Held-out cases | None during the run | Evaluation maintainers | Separate storage and credentials |
| Scorer | Execute through an interface | Evaluation maintainers | Pinned version |
| Pass rule | Read | Policy owner | Reviewed change process |
| Proposal output | Write | Agent | Treated as untrusted input |
| Verdict record | Read | Evaluator | Evaluator signs or writes it |
Trace effective permissions rather than repository layout. Two directories in one checkout are not separate if the agent can edit both. A protected branch is not separate if the agent holds credentials that can bypass its rules.
3. Put the evaluator under another authority
Choose a control the agent cannot grant to itself. Depending on the workflow, that could be:
- a separate repository that the agent identity can read but not write;
- an evaluation service with a narrow submission endpoint;
- a protected CI job that checks out the scorer at a pinned revision;
- read-only object storage for the corpus, with write access held by another identity.
Keep the submission channel narrow. The agent should submit the artifact being judged and the minimum metadata needed to reproduce it. It should not submit a replacement scorer, a new threshold, or an instruction that changes how the evaluator runs.
submission:
proposal_id: p-1842
baseline_id: b-1831
artifact_digest: sha256:...
task_contract: support-refund-v3
proposer_version: agent-2026-08-13
The evaluator resolves its own corpus, scorer, and acceptance rule from the declared task contract.
4. Build a corpus that can disagree
Use cases that separate a plausible proposal from an acceptable one. Include:
- ordinary cases that represent the intended task;
- edge cases near a policy or data boundary;
- regressions taken from earlier failures;
- counterexamples where a fluent or syntactically valid answer is still wrong;
- malformed inputs that should produce a safe refusal or error.
Keep held-out cases hidden from the agent during the run. Public examples can teach the task. Held-out examples check whether the resulting behavior extends beyond those examples.
Give each case a stable identifier. Add a reason when a case is quarantined, and keep the case visible in evaluation history. Silent deletion makes score changes difficult to interpret.
5. Make scoring reproducible
Pin every input that affects the verdict:
evaluation_id = hash(
corpus_version,
scorer_version,
rubric_version,
pass_rule_version,
runtime_version
)
Record those versions with the verdict. If scoring uses a model, record the model identifier and the exact evaluation prompt. If scoring uses a person, record the rubric version and the reviewer’s decision, but avoid storing personal data that the decision does not need.
Prefer explicit outcomes over one blended number. A proposal that improves 20 easy cases and breaks one protected case should not hide the regression inside an average.
6. Compare with the current baseline
Run the baseline and proposal against the same protected corpus. Classify each case:
IMPROVED proposal passes, baseline fails
UNCHANGED proposal and baseline have the same accepted result
REGRESSED baseline passes, proposal fails
UNCLEAR the scorer cannot distinguish them
Write the acceptance rule in terms of those classes. For example:
accept when:
required_checks_completed
and regressions == 0
and invalid_results == 0
and improved_cases >= 1
Change the rule through the evaluation owner’s process, not as part of the proposal under test.
7. Test the separation
Verify the boundary with attempted violations as well as a successful run.
- Ask the agent to edit a corpus case. Confirm the write is denied.
- Ask it to replace the scorer or lower the threshold. Confirm the evaluator ignores proposal-supplied evaluation assets.
- Remove evaluator access to one required input. Confirm the result is
NO_VERDICTand promotion stops. - Submit an artifact whose digest does not match its manifest. Confirm the
evaluator rejects it as
INVALID. - Run the same pinned submission twice. Confirm the recorded inputs and verdict are the same, or that any nondeterminism is reported.
- Add a known regression to a test proposal. Confirm the protected case catches it before using the evaluation for real decisions.
Common failure modes
The tests live elsewhere, but the credentials do not
Moving files does not help if the agent can still write through a shared token. Separate the writer identity and inspect effective permissions.
The agent can choose which checks run
A command such as evaluate --skip-held-out gives the proposal control over
the verdict. The evaluator should select required checks from the task contract.
Only the latest score is stored
A score without corpus, scorer, baseline, and runtime versions cannot explain why the result changed. Store the full evaluation identity with each verdict.
Evaluator downtime becomes success
Retries are reasonable. Promotion after exhausted retries is not. Record
NO_VERDICT, preserve the proposal, and wait for the evaluator to recover.
The corpus rewards the current agent
Add cases from production failures, human disagreements, and boundary conditions. Do not build the corpus only from examples the current agent already handles.
Related reading
- agent-systems
- evals
- trust-boundary
Keep exploring