> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Run Claude Code evals

> Connect Claude once, check one trial, then compare models

Sphynx runs the real Claude Code CLI in a sandbox. You don't install Claude locally or write an authentication script.

## Connect Claude

1. Select the organization that will own the evals.
2. Open **Settings > Harnesses > Add harness** and choose **Claude Code**.
3. Enter an Anthropic API key and choose **Everyone in the organization**, so CLI, SDK and CI batches can use it.
4. Save it. From the connection's actions menu, choose **Make default** if it isn't already, then **Check it works**.

The key stays in Sphynx. A local Claude login or Claude subscription does not authenticate these trials: Sphynx uses [Claude Code's bare mode](https://code.claude.com/docs/en/headless#start-faster-with-bare-mode), which requires API credentials.

Create a Sphynx API key with `evals:read` and `evals:write` in the same organization and set it as `SPHYNX_API_KEY` in your shell or CI secrets. Your eval file and GitHub workflow hold no Anthropic key. Hosted E2B sandboxes need no key of their own.

## Check one trial

Install `sphynx-sh`, then create `claude.eval.ts`:

```ts theme={null}
import { command, suite } from "sphynx-sh";

export default suite({
  id: "claude-smoke",
  prompt: "Create hello.txt containing exactly hello.",
  cases: [
    {
      id: "writes-hello",
      validate: command('test "$(cat hello.txt)" = hello'),
    },
  ],
  variants: [{ harness: "claude", model: "claude-haiku-4-5-20251001" }],
  trials: 1,
});
```

```sh theme={null}
npx sphynx-sh eval ./claude.eval.ts
```

This starts a batch with one run: the one case on the one variant, with one trial. A passing trial confirms the connection, the Claude install, its file tools and your check. Open the printed batch link to see the commands and output. One trial says nothing about how reliably the case passes.

## Compare models

Keep the cases and add variants:

```ts theme={null}
variants: [
  { harness: "claude", model: "claude-haiku-4-5-20251001" },
  { harness: "claude", model: "claude-sonnet-5" },
  { harness: "claude", model: "claude-opus-5" },
]
```

The next batch holds one run per case per variant, so the models appear side by side. Use [exact model IDs](https://platform.claude.com/docs/en/models/overview) your Anthropic account can call. Calls are billed to that account.

To test tool use, add [MCP, CLI or HTTP mocks](/evals/mocks). Sphynx sets them up in each sandbox for Claude the same way as for other harnesses.

A batch is capped at 100 trials (cases x variants x trials). Split a larger grid into separate suites. See [API limits](/api-reference/introduction#limits), and [CI](/guides/ci) for trial counts and the GitHub Action.

## If setup fails

* **Missing Claude credential:** check the organization, that the connection is shared with everyone, and that it is the default.
* **Authentication rejected:** rotate the Anthropic key and check it again.
* **Model unavailable:** check the exact ID and your account's access.
* **Works locally but fails in CI:** an API key uses the organization's connections, not your personal connection or local login.

Compiling the suite or passing local mock tests does not prove a hosted Claude trial works. Check one trial's result before growing the grid.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.