Skip to main content
Install sphynx-sh in your project and commit the lockfile. Name suite files *.eval.ts so the action can find them. Create an API key with evals:read and evals:write and save it as the SPHYNX_API_KEY repository secret. Use a dedicated organization for CI. In that organization, connect the harness under Settings > Harnesses and make it the default, shared with everyone in the organization. API keys cannot use personal connections. Provider credentials stay in Sphynx, not in GitHub or the suite.
Replace COMMIT_SHA with a reviewed action commit. The action runs the sphynx CLI you installed, with Node 20 or later. It does not install dependencies. Each suite file starts one batch: one run per case per variant. The job summary links each batch as soon as it starts. The dashboard labels these batches CI and links back to the GitHub Actions attempt, read from GitHub’s environment with no extra setup.
  • failures fails the job when any run fails or has a trial that did not pass.
  • strict also fails when a run or trial is missing or unfinished.
  • never ignores trial results.
Under every gate, a batch that fails to start, fails, or times out fails the job. The step exits 0 when the gate passes, 2 when it fails, and 1 when a batch could not start or finish. To keep the report, upload it with actions/upload-artifact and if: always().

What runs

A case runs in an empty workspace unless it or its suite names a source, such as repo("acme/api@8f31c4a"). The CLI does not point it at the checked-out commit for you. The checkout above uses the PR head, not GitHub’s merge commit, so the suites and validators are the ones in the PR. This example skips fork and Dependabot PRs because they don’t receive repository secrets. Keep unit tests in a separate job so they still run. Don’t use pull_request_target to run untrusted suites with a secret. The job is the PR check. For a separate check named sphynx on the PR head commit, pass github-token: ${{ github.token }} and grant the job checks: write. Its verdict follows the selected gate. Outside the action, the same command is:

Keep comparisons stable

When you test one change, keep the rest of the variant the same so its history stays comparable. Pin fixture repositories to a commit. A new harness release stays on the same variant, and each run records the harness version it used. Use at least three trials for a gating check. One trial shows an outcome, not whether it repeats.

Control the batch size

Five cases, two variants and three trials open 30 sandboxes. Keep pull request grids small and run larger comparisons on a schedule. A run holds 1 to 10 trials and a batch at most 100. See API limits. The CLI records batch IDs before waiting and never retries a start. If CI times out or is cancelled, the batch keeps running in Sphynx. Open its link before starting another.

What a trial costs

An agent re-sends its whole conversation every turn, so anything it reads early is paid for again on every later turn. Loading a large skill, or reading a dependency to work out its API, is the usual reason one trial costs ten times another with the same number of turns. Give the agent the interface instead of letting it search. A short reference in the workspace costs less than the files it would otherwise read, and keeps cost steadier: the trials that explore are the ones whose cost varies most. Each trial reports its token counts. The dashboard flags two things to act on: context that grew fast, and nothing served from cache. A grid multiplies both, so 30 sandboxes each carrying a large context cost 30 times as much.