> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Variants

> The harness, model and sandbox a case runs on

A variant is the harness, model and sandbox (plus an optional [profile](/evals/profiles)) a case runs on. A case can run on many variants, and comparing them, for example Codex against Claude Code, is the main reason to have more than one.

```ts theme={null}
{ harness: "codex", model: "gpt-5.6-sol" }
```

| Field | Purpose |
| - | - |
| `harness` | The agent that does the work |
| `model` | The model the harness uses |
| `sandbox` | Where each trial runs. Optional, defaults to `e2b` |
| `profile` | Optional `{ name, dir }`: files and settings layered on the harness. See [profiles](/evals/profiles) |

## Harnesses

| Agent | `harness` | Credential | Reports |
| - | - | - | - |
| Codex | `codex` | API key or ChatGPT | Commands, files, tokens |
| Claude Code | `claude` | API key | Commands, files, tokens |
| OpenCode | `opencode` | Auth file | Commands, files, tokens |
| Pi | `pi` | Auth file | Commands, files, tokens |
| Gemini CLI | `gemini` | API key | Commands, files, tokens |
| Qwen Code | `qwen` | API key | Commands, files, tokens |
| Cursor Agent | `cursor` | API key | Commands, files |
| FX | `fx` | AI Gateway key or ChatGPT auth file | Tool names, answer |

Each harness needs a [connection](/connections). Harnesses that report less are scored the same way but leave less trajectory and cost data. To run your own agent, use the [command harness](/evals/command-harness).

For Claude Code, follow the [setup guide](/guides/claude-code) and check one trial before comparing models.

## Models

Model ids differ by harness. List the ones a harness accepts:

```ts theme={null}
const { models } = await sphynx.evals.models.list({ harness: "codex" });
```

You don't have to list models before starting a batch. Lists come from Sphynx's Codex cache, models.dev or a fixed list, depending on the harness, so an empty result means the source is unavailable, not that the harness has no models.

## Sandboxes

A variant without `sandbox` runs on `e2b`. The others are `daytona`, `upstash`, `modal`, `cloudflare` and `vercel`. Every trial gets its own isolated workspace.

Hosted sandboxes can't reach services on your machine. To test against one, run the suite with `sphynx eval --local`. See [local runs](/evals/local).

The sandbox is part of the variant, so changing it starts a new history. If a variant ran on a sandbox other than `e2b`, keep naming it to keep adding to that history.

## Compare variants

Change one field at a time:

```ts theme={null}
variants: [
  { harness: "codex", model: "gpt-5.6-sol" },
  { harness: "codex", model: "gpt-5.6-terra" },
]
```

Starting the suite creates a batch with one run per case per variant, so the variants sit side by side.

A case has one variant for each combination of harness, model, sandbox, profile name and simulated-user model it has run on. Changing any of them creates a new variant with its own history. A new harness version stays in the same history, and each run records the version it used. Use a [profile](/evals/profiles) to compare an agent's stock and configured setups.

## Pick variants from the command line

Every variant in a suite file has a label. It is `harness/model`, plus `@` and the profile name when it has a profile. `sphynx eval` prints these labels, and `--variant` takes them back:

```ts theme={null}
variants: [
  { harness: "codex", model: "luna", profile: { name: "autumn-setup", dir: "./autumn" } },
  { harness: "codex", model: "luna" },
]
```

| `--variant` | Runs |
| - | - |
| `codex/luna` | The variant with no profile |
| `codex/luna@autumn-setup` | The variant with the profile |
| `autumn-setup` | The variant with the profile, named by the profile alone |

A label matches its own variant first. `harness/model` also picks the one variant on that model with a profile, when no variant without one exists. When a name matches no variant, or more than one, the command stops and lists the labels to choose from. Repeat `--variant` to run more than one.

## Read a variant's history

```ts theme={null}
const { runs, total } = await sphynx.evals.runs.list({ caseId: "writes-hello" });
```

Runs come newest first, 20 to a page. Pass `variant` (a variant id) to read one variant, and `page` for older runs. `evals.cases.get({ id })` lists a case's variants with their ids.

To run a case again, `evals.cases.run({ id, trials, variants })` starts a batch that runs the case's newest version on the variant ids you name, or on every variant it has run on when you leave `variants` out. `trials` defaults to 1.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.