sphynx eval import converts an existing case file into a *.eval.ts suite.
--out, the suite prints to stdout. A summary of converted and unconverted assertions goes to stderr.
The generated suite has a placeholder variant (
codex on gpt-5.6-sol) and trials: 3, with each case’s text passed through a task variable.
evals-json
skill_name and evals are required, and each case needs id, prompt and assertions. An assertion is a structured object or a prose string.
Matching is case-insensitive. Prose assertions cannot be converted safely, so each becomes a placeholder that always fails:
expected_output becomes a comment, not a check. files names directories but not their contents, so each case starts with files({}) and a comment naming them.
yaml
A directory imports every*.yaml and *.yml file in it, sorted by name.
name and task are required. The task becomes the prompt. Each judge_context entry needs human judgment, so it becomes a failing placeholder. A case without judge_context gets one placeholder asking what a good answer is.
max_steps becomes a comment, because Sphynx does not cap agent steps. YAML cases start with files({}) because the format carries no source files.
Finish the suite
- Review the generated file.
- Replace each
files({})with fixture files or a repository source. - Set the harness, model and sandbox for each variant.
- Replace every
unwritten(...)placeholder with a real check. - Run it with
sphynx eval ./suite.eval.ts.