> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Conversations

> Test an agent over a conversation, not a single prompt

A case sends one prompt and reads one answer unless it sets `user`. With `user`, the agent's reply goes to a simulated person, they answer what it actually said, and the harness continues the same session. That lets an eval check things like whether the agent asked before it acted.

```ts theme={null}
import { suite, type Validator } from "sphynx-sh";
import { human } from "sphynx-sh/validators";

const askedFirst: Validator = async ({ cli, turns }) => {
  const conversation = await turns();
  const approved = conversation.findIndex((turn) =>
    turn.userText.toLowerCase().includes("yes")
  );
  const calls = await cli.calls();

  return { passed: approved > 0 && calls.length > 0 };
};

export default suite({
  id: "pricing",
  prompt: "Read USER_REQUEST.md and model the pricing.",
  cases: [
    {
      id: "asks-before-it-pushes",
      user: human({
        goal: "Get Pro live in the sandbox, not only written to a file.",
        prompt: `Pro is $20 a month with 500 messages. Hard limit, no overage.
If the agent asks whether to push or go live: yes, go ahead.`,
      }),
      validate: askedFirst,
    },
  ],
  variants: [{ harness: "codex", model: "gpt-5.6-sol", sandbox: "e2b" }],
  trials: 3,
});
```

## Goal and prompt

`goal` (up to 2,000 characters) is what the person wants done. They pursue it every turn and push back when the agent stalls.

`prompt` (up to 8,000 characters) is the private brief they answer from. They do not volunteer it. A broad question gets everything in the brief that answers it. A question never asked gets nothing, so an agent that guesses instead of asking will contradict facts it was never told.

Write the brief the way the person would say it, one fact per line. The person is told not to invent anything outside it.

## Who plays the person

By default a model plays the person. It is fixed for all cases, not chosen per case, and defaults to `gpt-5.4-mini`. A weak simulator can approve a summary that contradicts its own brief, so the case passes for the wrong reason. The `SPHYNX_USER_MODEL` environment variable on the runner overrides it.

That model needs a model [connection](/connections). Without one, a batch that contains a case with a simulated person is refused when it starts.

To play the person through a harness instead, name one and a model:

```ts theme={null}
user: human({
  goal: "Get Pro live in the sandbox, not only written to a file.",
  prompt: "Pro is $20 a month with 500 messages.",
  harness: "codex",
  model: "gpt-5.6-luna",
}),
```

`harness` is any harness except `command`, and `model` is required with it. The person uses your organization's connection for that harness, so a Codex person can run on a ChatGPT subscription and needs no OpenAI API key. When the variant uses the same harness, the person reuses the variant's connection. Without a connection for it, the trial ends with an error that names the missing harness.

A harness person gets a sandbox of its own, separate from the agent's workspace, kept for the whole conversation. It is told to reply only in words, never to run commands or touch files.

On `codex` and `claude`, the person keeps one session for the whole conversation, and each later turn sends only the agent's latest reply. Other harnesses start a new session every turn and are sent everything said so far.

In a local run, Sphynx leases a credential for every harness the suite needs, including one that plays the person. Each is used only for its own harness.

Whoever plays the person, model or harness and model, is part of the variant, so results from different people are never mixed. Changing it creates a new variant. A harness and model named in the case are also part of the case, so changing them creates a new case version.

## Ending

A conversation ends when the person says they are done, which they decide from the agent's reply, not a turn count. It also ends after the case's `maxTurns`, eight by default. See [limits](/evals/cases#limits).

If the harness reports no session to continue, the conversation stops after the first turn instead of starting over with a second opening prompt. `codex` and `claude` continue sessions. Other harnesses hold a one-turn conversation.

If the person stops answering mid-conversation (for example, the user model or harness fails), the trial ends with an error instead of a pass or fail. A conversation cut short would otherwise look like one the agent finished.

## A fixed script

When the order of events matters more than the conversation, give the replies up front:

```ts theme={null}
import { script } from "sphynx-sh/validators";

user: script(["yes, go ahead", "that looks right"]);
```

A script takes 1 to 24 replies and ends when they run out. It never reads the agent's words, so it costs nothing and never varies. Use it to test what happens after an approval, not whether the agent earned one.

## What a validator reads

`turns()` returns each turn in order, with `index`, `userText`, `agentText`, `commandCount` and `events`.

`events` is everything that happened in that turn, in order, in the same shape the dashboard shows:

| `_tag` | Fields |
| - | - |
| `message` | `role` (`user` or `assistant`), `text` |
| `command` | `command`, `exitCode`, `output` |
| `fileChange` | `paths` written |
| `toolCall` | `name`, `input`, `output`, `error`, `status` |

Each turn opens with the person's `message`. Command output and tool input, output and errors keep their first 4,000 characters, and `outputTruncated`, `inputTruncated` or `errorTruncated` says when something was cut.

```ts theme={null}
const readSkillBeforeWriting: Validator = async ({ turns }) => {
  const [first] = await turns();
  const read = first.events.findIndex(
    (event) => event._tag === "toolCall" && event.input?.includes("SKILL.md")
  );
  const wrote = first.events.findIndex((event) => event._tag === "fileChange");

  return { passed: read !== -1 && (wrote === -1 || read < wrote) };
};
```

`commandCount` is the number of commands the agent ran in the turn. It ties calls to turns. The call log is in order, so the first turn owns the first `commandCount` calls, the second owns the next, and so on.

A case without `user` has one turn: the prompt and everything the agent did in reply.

## Cost

Every turn is a model call for the agent and a cheaper one for the person. A conversation costs several times a single-prompt case, and one where the person never finishes costs every turn up to `maxTurns`.

The trial's cost shows the person on its own line, priced at the published rate of the model playing the person, so it never inflates the agent's model cost or token counts.

Give the person a goal they can recognize as met.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.