> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Run evals in CI

> Gate pull requests on eval results

Install `sphynx-sh` in your project and commit the lockfile. Name suite files `*.eval.ts` so the action can find them.

Create an API key with `evals:read` and `evals:write` and save it as the `SPHYNX_API_KEY` repository secret. Use a dedicated organization for CI.

In that organization, connect the harness under **Settings > Harnesses** and make it the default, shared with everyone in the organization. API keys cannot use personal connections. Provider credentials stay in Sphynx, not in GitHub or the suite.

```yaml theme={null}
name: Agent evals

on:
  pull_request:
    branches: [main]

permissions:
  contents: read

jobs:
  eval:
    if: >-
      github.actor != 'dependabot[bot]' &&
      github.event.pull_request.head.repo.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v6
        with:
          ref: ${{ github.event.pull_request.head.sha }}
          persist-credentials: false
      - uses: oven-sh/setup-bun@v2
      - run: bun install --frozen-lockfile
      - uses: charlietlamb/sphynx/.github/actions/eval@COMMIT_SHA
        with:
          api-key: ${{ secrets.SPHYNX_API_KEY }}
```

Replace `COMMIT_SHA` with a reviewed action commit. The action runs the `sphynx` CLI you installed, with Node 20 or later. It does not install dependencies.

Each suite file starts one batch: one run per case per variant. The job summary links each batch as soon as it starts. The dashboard labels these batches **CI** and links back to the GitHub Actions attempt, read from GitHub's environment with no extra setup.

| Input | Default | Purpose |
| - | - | - |
| `api-key` | Required | Sphynx API key |
| `working-directory` | `.` | Project with `sphynx-sh` installed |
| `file` | Discover `*.eval.ts` | Run one suite file |
| `case` | Every case | Run one case. With `file`, picks it from the file. Without, runs the stored case again |
| `variant` | Every variant | Variants to run, separated by commas or new lines: `harness/model`, `harness/model@profile` or a profile name for a file, variant ids for a stored case |
| `fail-on` | `failures` | `failures`, `strict` or `never` |
| `timeout` | `1200` | Seconds to wait per batch |
| `github-token` | None | Token that posts the results as a check run. Needs `checks: write` |

* `failures` fails the job when any run fails or has a trial that did not pass.
* `strict` also fails when a run or trial is missing or unfinished.
* `never` ignores trial results.

Under every gate, a batch that fails to start, fails, or times out fails the job. The step exits 0 when the gate passes, 2 when it fails, and 1 when a batch could not start or finish.

| Output | Value |
| - | - |
| `report` | Path to a JSON file with each suite file's batch ID, batch and problems |
| `batch-ids` | The batch IDs started, separated by commas |
| `conclusion` | `success`, `failure`, or `neutral` when a batch was started without waiting |

To keep the report, upload it with `actions/upload-artifact` and `if: always()`.

## What runs

A case runs in an empty workspace unless it or its suite names a [`source`](/evals/cases#sources), such as `repo("acme/api@8f31c4a")`. The CLI does not point it at the checked-out commit for you. The checkout above uses the PR head, not GitHub's merge commit, so the suites and validators are the ones in the PR.

This example skips fork and Dependabot PRs because they don't receive repository secrets. Keep unit tests in a separate job so they still run. Don't use `pull_request_target` to run untrusted suites with a secret.

The job is the PR check. For a separate check named `sphynx` on the PR head commit, pass `github-token: ${{ github.token }}` and grant the job `checks: write`. Its verdict follows the selected gate.

Outside the action, the same command is:

```sh theme={null}
bunx --no-install sphynx eval --fail-on failures --timeout 1200 --output results.json
```

## Keep comparisons stable

When you test one change, keep the rest of the [variant](/evals/variants) the same so its history stays comparable. Pin fixture repositories to a commit. A new harness release stays on the same variant, and each run records the harness version it used.

Use at least three trials for a gating check. One trial shows an outcome, not whether it repeats.

## Control the batch size

```text theme={null}
trials in a batch = cases x variants x trials per run
```

Five cases, two variants and three trials open 30 sandboxes. Keep pull request grids small and run larger comparisons on a schedule.

A run holds 1 to 10 trials and a batch at most 100. See [API limits](/api-reference/introduction#limits).

The CLI records batch IDs before waiting and never retries a start. If CI times out or is cancelled, the batch keeps running in Sphynx. Open its link before starting another.

## What a trial costs

An agent re-sends its whole conversation every turn, so anything it reads early is paid for again on every later turn. Loading a large skill, or reading a dependency to work out its API, is the usual reason one trial costs ten times another with the same number of turns.

Give the agent the interface instead of letting it search. A short reference in the workspace costs less than the files it would otherwise read, and keeps cost steadier: the trials that explore are the ones whose cost varies most.

Each trial reports its token counts. The dashboard flags two things to act on: context that grew fast, and nothing served from cache. A grid multiplies both, so 30 sandboxes each carrying a large context cost 30 times as much.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.