Skip to main content
Starting evals creates a batch with one run per case per variant. Each run holds its trials. evals.batches.start returns the batch id and, for each case and variant, a run with its caseId and variantId, while trials continue on the server. startAndWait resolves once the batch stops:
To wait on a batch you started earlier, pass its id to evals.batches.wait({ id }). Both poll evals.batches.get, backing off as the batch goes on, and both take onProgress, signal, timeoutMs, pollIntervalMs and maxIntervalMs. Neither cancels the batch. On timeout or abort it keeps going, and evals.batches.get can still read it. Each run in batch.runs names its case, suite and variant, and has its own status, trials and distribution.

Run distribution

A run’s pass rate is taken across its trials:
Only scored trials count. When scored is zero, passRate is zero, which is not a result. deterministic is true when at least two scored trials all agree on pass or fail and their command counts are within four of each other.

Trial status

A void trial is an infrastructure failure, not a verdict. voidFields names the missing evidence where it can be identified. When a trial stopped before it could be scored, failure says why, such as The agent ran past its time limit of 5m or a setup step that failed; it is null for a trial that was scored. The dashboard shows the same reason on the trial. Keep void trials out of pass rates in reports and CI. Coding trials take minutes, so a batch that stays running is usually still working. A server restart does not resume active work, and runs abandoned for six hours are marked failed. A local batch from a CLI that checks in is marked failed once its machine has been quiet for 10 minutes.

Trajectory

Each trial has an ordered trajectory of commands, messages, tool calls and file changes:
Trials also record command counts, changed files, model and sandbox time, the sandbox id, and token counts when the harness reports them. Command output and tool payloads are capped at 4,000 characters, and truncation is marked. Usage carries a warning when a trial re-read a large context on every turn or got nothing from cache. Both mean the trial paid again and again for something it read once. See what a trial costs. Start with the distribution, then compare a passing and a failing trajectory from the same run. The pass rate tells you what happened. The command counts tell you how the attempts differed.

Validation evidence

Each trial’s validations holds the code checks, judges and command that scored it:
Code checks record their return value, context calls and results, logs and exceptions. Judges record the prompt, raw response, provider metadata and parsed judgment. Commands record the command, stdout, stderr and exit code. failed means a check finished and did not pass, error means it could not finish, and skipped means it never ran. Every code check runs even when an earlier one fails or crashes, so each keeps its own result. A check is skipped only when the trial ended before it could start, or when it is a judge and the trial was already void. Runs from before capture existed have no evidence to show. Each validator gets a 64,000-character budget. Single values cap at 16,000 characters, and at most 64 calls or log entries are kept. Source snapshots cap at 100 files and 1,000,000 characters, and older runs without one show “Source unavailable.” Truncation affects stored evidence, never the score. Inputs and logs can contain your data. Set captureValidation: false on the suite to keep verdicts and timing but store no payloads. captureSource controls source snapshots separately. Model reasoning is never collected.

Secrets in evidence

Sphynx does not store the value your prepare returns. A check can read prepared and pass or fail on it, and its evidence shows no input. Before evidence is stored or shown, every value, log line and message is checked for credentials. Anything that looks like a key is replaced with [redacted]. That covers keys starting with sk-, sk_live_, am_sk_, anp_, ghp_, github_pat_, xox, AKIA and AIza, JSON web tokens, private key blocks, the token after Bearer, and the password in a URL such as postgres://app:password@db. Every value of the harness and sandbox credentials the trial ran with is replaced too, even when a check reads it back from a file in the sandbox. A forwarded host variable is replaced when its name says it is secret, such as STRIPE_SECRET_KEY, or when its value looks like a token. A forwarded address such as http://localhost:3000 stays readable, and only a password or token inside it is replaced. The agent’s journal gets the same treatment before it is stored, streamed to the dashboard or printed in your terminal. That covers commands, their output, tool calls, messages and changed file paths. Failure reasons, such as the error output of a prepare or a clone, are redacted too. Code checks still read what the agent actually wrote, so redaction never changes their verdict. Judges read the redacted conversation, so a judge model is never sent a credential Sphynx recognizes. The files a trial saves from its workspace are redacted the same way before they are stored. Only small text files are saved, so a binary file is never saved or changed. A file whose redacted text still holds something that looks like a password assignment is not saved at all. Text is redacted before it is cut to the limits above, so a key that crosses a limit never leaves a fragment behind. A secret with no known shape, such as a plain password, is stored if a check or the agent prints it. Log a summary instead of the value. Evidence stored before redaction existed is kept as it was. If an older run recorded a key, rotate that key.

Trigger

Each batch records a trigger (Dashboard, CLI, CI, API or MCP), and its runs inherit it. Sphynx sets it for you. The CLI records CLI, or CI with a link to the GitHub Actions attempt when it runs there. Batches started through the API record API. The trigger describes the caller. It grants and verifies nothing. Running a case again records the new trigger, and older runs show Unknown.