> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Quickstart

> Define a suite and run it on a coding agent

`sphynx-sh` is the TypeScript package for defining and running evals for coding agents. Suites, mocks and validators live in your repository next to the code they test.

## Prerequisites

Add a [harness connection](/connections), then create an API key with `evals:write` and `evals:read` under **Settings > API keys**.

```bash theme={null}
export SPHYNX_API_KEY="anp_..."
```

Save harness credentials in Sphynx, never in eval definitions. Hosted sandboxes and mocks need no keys of their own.

## Install

<CodeGroup>
  ```bash npm theme={null}
  npm install sphynx-sh
  ```

  ```bash bun theme={null}
  bun add sphynx-sh
  ```
</CodeGroup>

## Define a suite

Create `smoke.eval.ts`:

```ts theme={null}
import { command, suite } from "sphynx-sh";

export default suite({
  id: "smoke",
  prompt: "{{task}}",
  cases: [
    {
      id: "writes-hello",
      variables: { task: "Create hello.txt containing exactly hello" },
      validate: command('test "$(cat hello.txt)" = hello'),
    },
  ],
  variants: [{ harness: "codex", model: "gpt-5.6-sol" }],
  trials: 3,
});
```

## Run it

```bash theme={null}
npx sphynx-sh eval ./smoke.eval.ts
```

The CLI compiles the file on your machine, including imported fixtures and validators, and starts a batch. Each trial runs in its own hosted sandbox. The [GitHub Action](/guides/ci) runs the same command, and you can also [start a batch from your own code](/guides/run-from-code).

## Vocabulary

* **Suite:** a system prompt and shared setup for many cases. It needs an `id`, and its `name` defaults to the id.
* **Case:** one eval. It needs an `id`, and its `name` defaults to the id. It sets up a scenario and checks the result with `validate`. A case gets a new version when its definition changes.
* **Variant:** the harness, model and sandbox, plus an optional profile, a case runs on.
* **Run:** one case on one variant. It holds the trials.
* **Trial:** one attempt, in its own sandbox. A run's pass rate is taken across its trials.
* **Batch:** the runs started together. Starting evals creates a batch.

The suite above has one case, one variant and three trials, so its batch holds one run with three trials. Two cases on three variants with three trials each would be six runs and 18 trials.

## Package exports

| Import | Use |
| - | - |
| `sphynx-sh` | [`suite`, `command`, `repo`, `files`, `empty` and validator types](/evals/cases) |
| `sphynx-sh` | [`Sphynx` client for batches, cases and runs](/sdk/client) |
| `sphynx-sh/mcp` | [`server`, `tool`, `resource`](/evals/mocks#mcp) |
| `sphynx-sh/cli` | [`cli`, `command`](/evals/mocks#cli) |
| `sphynx-sh/api` | [`api`, `endpoint`, `withApi`](/evals/mocks#http-apis) |
| `sphynx-sh/validators` | [`judge`](/evals/judges), [`human`, `script`](/evals/conversations) |
| `sphynx-sh/eval` | [`compileEval`, `compileDefinition`](/guides/run-from-code) |

Definitions are plain objects and handlers are ordinary functions. Mock inputs use Standard Schema for types and runtime checks (the examples use Zod). You don't need Effect or a separate model client.

<CardGroup cols={2}>
  <Card title="Write cases" icon="list-check" href="/evals/cases">
    Add sources, preparation and typed validators.
  </Card>

  <Card title="Mock dependencies" icon="terminal" href="/evals/mocks">
    Give agents typed MCP servers, CLIs and HTTP APIs.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.