> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sphynx.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Command harness

> Run your own agent as a process inside the sandbox

The `command` harness runs a process you own instead of a coding agent Sphynx installs. Sphynx starts it in the sandbox, passes it the case through environment variables, reads events from its stdout, and scores the trial with the case's `validate`.

```ts theme={null}
variants: [
  {
    harness: "command",
    model: "openai/gpt-5.5",
    profile: { name: "sample", dir: "./profile" },
    sandbox: "e2b",
  },
]
```

The command harness needs a [profile](/evals/profiles). Its `run` is the command line, run with `bash -c` from the workspace. Its `install` runs once while the sandbox is prepared. The profile directory ships everything else the process needs.

## Environment

| Variable | Value |
| - | - |
| `SPHYNX_PROMPT` | The case prompt |
| `SPHYNX_MODEL` | The variant's `model`, verbatim |
| `SPHYNX_HOME` | The sandbox home directory |
| `SPHYNX_WORKSPACE` | The workspace, and the working directory when the process starts |
| `SPHYNX_SYSTEM_PROMPT_FILE` | Path to the profile's system prompt, when it has one |
| `SPHYNX_TRACE_LOG` | Where the shell recorder appends what it sees |

The profile's `env` is set too. Stdin is closed.

## Runner kit

Write the process in TypeScript with `sphynx-sh/runner`. It reads the environment and prints the event lines for you, and it rejects a line the server would not accept.

```ts theme={null}
import { createEmitter, env } from "sphynx-sh/runner";

const { model, prompt } = env();
const emit = createEmitter();

emit.started({ model, sessionId: "run-1" });
emit.toolCall({ name: "search", input: { query: prompt }, output: ["total.ts"] });
emit.message({ text: "The total lives in total.ts." });
emit.usage({ inputTokens: 900, outputTokens: 40 });
emit.finished("done");
```

Set the profile's `run` to start the script, for example `"run": "bun run agent.ts"`.

`env()` returns `prompt`, `model`, `workspace`, `home`, `systemPromptFile` and `traceLog`. It throws when a variable the harness always sets is missing, so a script run by hand fails with the name of the variable.

| Method | Line it prints |
| - | - |
| `started({ sessionId, model })` | `Started` |
| `message({ text, role?, usage? })` | `Message`. `role` defaults to `assistant` |
| `toolCall({ name, input, output?, error?, status?, callId? })` | `ToolCall`. A value that is not a string is written as JSON |
| `usage({ inputTokens, outputTokens, cacheReadTokens?, cacheWriteTokens? })` | `Usage`. Counts must be whole numbers, zero or more |
| `command({ command, exitCode, output })` | `Command` |
| `fileChange(paths)` | `FileChange` |
| `finished(reason)` | `Finished`. A second call prints nothing |

Every event carries the time it was printed. To collect lines instead of printing them, pass `createEmitter({ write, now })`.

## Events

The kit prints the lines below. They are also the protocol for a process in any other language.

Print one JSON object per line on stdout. Lines with a known `_tag` are recorded and every other line is ignored, so ordinary logging is fine. `at` is epoch milliseconds and optional. A line without it is stamped when it is read.

```json theme={null}
{"_tag":"Started","sessionId":"run-42","model":"openai/gpt-5.5"}
{"_tag":"Message","role":"assistant","text":"Reading the failing test first."}
{"_tag":"Command","command":"bun test","exitCode":1,"output":"expect(5).toBe(6)\n"}
{"_tag":"FileChange","paths":["/workspace/src/total.ts"]}
{"_tag":"ToolCall","callId":"call_7","name":"search","input":"{\"query\":\"total\"}","status":"completed"}
{"_tag":"Usage","inputTokens":9120,"outputTokens":312,"cacheReadTokens":8000}
{"_tag":"Finished","reason":"done","at":1788300005000}
```

| Line | Recorded as |
| - | - |
| `Started` | The session id and model. Optional. Without it the trial has no session id |
| `Message` | A conversation turn. `role` is `assistant` or `user`. The last assistant message is the answer validators read |
| `Command` | A shell command with its exit code and output. `exitCode` is required but may be `null` |
| `FileChange` | Absolute paths written |
| `ToolCall` | A tool call by name. `callId` and `status` are required but may be `null` |
| `Usage` | Tokens for one turn, added to the trial's total. `cacheReadTokens` and `cacheWriteTokens` default to 0, `totalTokens` to input plus output |
| `Finished` | Ends the journal with a reason |

## Exit

A non-zero exit does not fail the trial or discard earlier events. If the process never prints `Finished`, Sphynx records one with the reason `exit N`. The trial is scored either way.

The process shares the case's [time limit](/evals/cases#limits) with every other turn of the trial, 15 minutes by default. If it is still running when the limit runs out, it is stopped and the trial ends as timed out.

## The recorder

Sphynx sources a small script into every non-interactive bash the process starts, through `BASH_ENV`. Its `DEBUG` trap appends one line per command to `SPHYNX_TRACE_LOG`. Each line becomes a `Command` in the journal with a `null` exit code, unless you already printed a `Command` with exactly the same text. The log is cleared after each turn, so a turn records only the commands it ran.

It only sees bash and zsh, the two shells with a `DEBUG` trap:

* `bash -c` and nested bash scripts are traced line by line.
* `sh -c` on Debian is dash, which has no `DEBUG` trap. Only the `sh -c …` invocation is recorded.
* Shells spawned by Node, Python or other runtimes are invisible. A Python agent leaves one line: its own invocation.
* Exit codes are never captured, because the trap fires before the command runs.

To get inner commands and exit codes into the journal, print `Command` lines yourself.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.