Harnesses
Each harness needs a connection. Harnesses that report less are scored the same way but leave less trajectory and cost data. To run your own agent, use the command harness.
For Claude Code, follow the setup guide and check one trial before comparing models.
Models
Model ids differ by harness. List the ones a harness accepts:Sandboxes
A variant withoutsandbox runs on e2b. The others are daytona, upstash, modal, cloudflare and vercel. Every trial gets its own isolated workspace.
Hosted sandboxes can’t reach services on your machine. To test against one, run the suite with sphynx eval --local. See local runs.
The sandbox is part of the variant, so changing it starts a new history. If a variant ran on a sandbox other than e2b, keep naming it to keep adding to that history.
Compare variants
Change one field at a time:Pick variants from the command line
Every variant in a suite file has a label. It isharness/model, plus @ and the profile name when it has a profile. sphynx eval prints these labels, and --variant takes them back:
A label matches its own variant first.
harness/model also picks the one variant on that model with a profile, when no variant without one exists. When a name matches no variant, or more than one, the command stops and lists the labels to choose from. Repeat --variant to run more than one.
Read a variant’s history
variant (a variant id) to read one variant, and page for older runs. evals.cases.get({ id }) lists a case’s variants with their ids.
To run a case again, evals.cases.run({ id, trials, variants }) starts a batch that runs the case’s newest version on the variant ids you name, or on every variant it has run on when you leave variants out. trials defaults to 1.