Combine checks
validate accepts one validator or judge, or an array of them:
Configuration
A case accepts up to 20 validators.
judge() checks its options when you call it.
Grade workspace files
Usefiles when the result is a file, not the final answer. A judge that grades a blog post can read the post itself:
...
If a file is missing or over a limit, the judge is not asked. Its judgment is invalid with an error that names the file, and the trial is void. It never counts as a pass or a fail.
Authentication
A harness judge uses your organization’s connection for that harness. A Codex judge can use a ChatGPT subscription connection and needs no OpenAI API key. When the variant uses the same harness, the judge reuses the variant’s connection. Aprovider: "openai" judge needs its own OpenAI credential. See connections.
Evidence and isolation
A judge receives the rendered case prompt, the conversation, the final answer andexpected, if set. The conversation lists what happened in order: each user and agent message, each command with its exit code and output, each tool call, and each file written. A case without user is a one-turn conversation that opens with the prompt, so a judge can still ask what the agent did before it answered.
Command output, tool input and output, and long messages are cut to 4,000 characters. A conversation longer than 48,000 characters keeps its start and end and marks how many steps were left out of the middle.
A judge does not see the rest of the workspace, MCP or CLI configuration, or sandbox environment. To grade a file the agent wrote, name it in files.
A harness judge starts in a fresh sandbox after the trial’s sandbox closes. It is told to grade only the supplied evidence. If it uses tools anyway, the judgment is invalid. Tool use is detected, not blocked by a firewall. To assert on tool use, keep a code validator that reads the mock call log.
An OpenAI judge uses structured outputs with storage disabled. Both kinds must pick one of the declared choices and give a short reason. A timeout, refusal, invalid response or provider failure gives score: null and voids the trial, so it is not counted as a pass or a fail.
Results
Each trial’s judgments carryname, model, evaluator, choice, score, threshold, reason, durationMs, error and the tokens the judge used as usage. The dashboard shows them apart from code checks.
Judging has its own line in the trial’s cost, apart from the agent’s model. Each judge is priced at its model’s published rate. If a judge reported no usage, or its model has no published rate, the line reads as not known and the trial’s cost is marked incomplete. The sandbox a harness judge runs in is not priced.
For the request, raw response and provider metadata, open the judge entry in the trial’s validations. Invalid response text is kept even when no score is assigned. See validation evidence for capture limits and privacy settings.
Changing a judge changes the case definition, so later runs record a new case version. Calibrate each judge prompt against known good and bad answers before trusting it, and use several trials: judgments vary even with the same model and prompt.