> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Model judges

> Score answers with a model you choose

Use code for exact checks. Add a judge for qualities that need interpretation, such as correctness or clarity.

```ts theme={null}
import { suite } from "sphynx-sh";
import { judge } from "sphynx-sh/validators";

const correctness = judge({
  name: "correctness",
  harness: "codex",
  model: "gpt-5.6-sol",
  prompt: "The answer explains why the result is correct, without inventing facts.",
  expected: "2 + 2 = 4",
  choices: { correct: 1, partial: 0.5, incorrect: 0 },
  threshold: 1,
});

export default suite({
  id: "arithmetic",
  prompt: "What is 2 + 2? Explain briefly.",
  cases: [{ id: "addition", validate: correctness }],
  variants: [{ harness: "codex", model: "gpt-5.6-sol", sandbox: "e2b" }],
  trials: 3,
});
```

The judge model is independent of the variant being measured. Sphynx requests exactly the model you name. If it is unavailable, the judgment is unscored. There is no fallback model.

## Combine checks

`validate` accepts one validator or judge, or an array of them:

```ts theme={null}
validate: [checkToolCalls, correctness];
```

Code checks run in the trial's sandbox. Judges run after it. A trial passes only when every code check passes and every judge meets its threshold. A passing judge cannot override a failed code check.

## Configuration

| Field | Meaning |
| - | - |
| `name` | Judge name, unique within the case |
| `harness` or `provider` | A harness (any except `command`), or `provider: "openai"` for a direct model call |
| `model` | Exact model identifier |
| `prompt` | Instructions for the judgment, separate from the case prompt |
| `expected` | Optional reference answer |
| `files` | Optional workspace files the judge reads, 1 to 8 paths relative to the workspace |
| `choices` | Labels mapped to scores from 0 to 1 (1 to 20 choices) |
| `threshold` | Inclusive pass threshold, default 1 |
| `timeoutMs` | Total judging timeout, default 120,000, range 1,000 to 300,000 |

A case accepts up to 20 validators. `judge()` checks its options when you call it.

## Grade workspace files

Use `files` when the result is a file, not the final answer. A judge that grades a blog post can read the post itself:

```ts theme={null}
const quality = judge({
  name: "post quality",
  harness: "codex",
  model: "gpt-5.6-sol",
  prompt: "The post is accurate, clear and ready to publish.",
  choices: { publish: 1, revise: 0 },
  files: ["out/post.md"],
});
```

Sphynx reads each file after the code checks run, before the trial's sandbox closes. The judge gets the full text of each file with its path. Files are not cut like the conversation. Secrets are redacted the same way as in the answer.

Each file can hold up to 32,000 characters. All judge files in a case together can hold up to 96,000. A path must be relative to the workspace and cannot contain `..`.

If a file is missing or over a limit, the judge is not asked. Its judgment is invalid with an error that names the file, and the trial is void. It never counts as a pass or a fail.

## Authentication

A harness judge uses your organization's connection for that harness. A Codex judge can use a ChatGPT subscription connection and needs no OpenAI API key. When the variant uses the same harness, the judge reuses the variant's connection.

A `provider: "openai"` judge needs its own OpenAI credential. See [connections](/connections).

## Evidence and isolation

A judge receives the rendered case prompt, the conversation, the final answer and `expected`, if set. The conversation lists what happened in order: each user and agent message, each command with its exit code and output, each tool call, and each file written. A case without `user` is a one-turn conversation that opens with the prompt, so a judge can still ask what the agent did before it answered.

Command output, tool input and output, and long messages are cut to 4,000 characters. A conversation longer than 48,000 characters keeps its start and end and marks how many steps were left out of the middle.

A judge does not see the rest of the workspace, MCP or CLI configuration, or sandbox environment. To grade a file the agent wrote, name it in [`files`](#grade-workspace-files).

A harness judge starts in a fresh sandbox after the trial's sandbox closes. It is told to grade only the supplied evidence. If it uses tools anyway, the judgment is invalid. Tool use is detected, not blocked by a firewall. To assert on tool use, keep a code validator that reads the mock call log.

An OpenAI judge uses structured outputs with storage disabled. Both kinds must pick one of the declared choices and give a short reason. A timeout, refusal, invalid response or provider failure gives `score: null` and voids the trial, so it is not counted as a pass or a fail.

## Results

Each trial's judgments carry `name`, `model`, `evaluator`, `choice`, `score`, `threshold`, `reason`, `durationMs`, `error` and the tokens the judge used as `usage`. The dashboard shows them apart from code checks.

Judging has its own line in the trial's cost, apart from the agent's model. Each judge is priced at its model's published rate. If a judge reported no usage, or its model has no published rate, the line reads as not known and the trial's cost is marked incomplete. The sandbox a harness judge runs in is not priced.

For the request, raw response and provider metadata, open the judge entry in the trial's `validations`. Invalid response text is kept even when no score is assigned. See [validation evidence](/evals/results#validation-evidence) for capture limits and privacy settings.

Changing a judge changes the case definition, so later runs record a new case version. Calibrate each judge prompt against known good and bad answers before trusting it, and use several trials: judgments vary even with the same model and prompt.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.