Skip to main content
sphynx eval import converts an existing case file into a *.eval.ts suite.
Without --out, the suite prints to stdout. A summary of converted and unconverted assertions goes to stderr. The generated suite has a placeholder variant (codex on gpt-5.6-sol) and trials: 3, with each case’s text passed through a task variable.

evals-json

skill_name and evals are required, and each case needs id, prompt and assertions. An assertion is a structured object or a prose string. Matching is case-insensitive. Prose assertions cannot be converted safely, so each becomes a placeholder that always fails:
expected_output becomes a comment, not a check. files names directories but not their contents, so each case starts with files({}) and a comment naming them.

yaml

A directory imports every *.yaml and *.yml file in it, sorted by name.
name and task are required. The task becomes the prompt. Each judge_context entry needs human judgment, so it becomes a failing placeholder. A case without judge_context gets one placeholder asking what a good answer is. max_steps becomes a comment, because Sphynx does not cap agent steps. YAML cases start with files({}) because the format carries no source files.

Finish the suite

  1. Review the generated file.
  2. Replace each files({}) with fixture files or a repository source.
  3. Set the harness, model and sandbox for each variant.
  4. Replace every unwritten(...) placeholder with a real check.
  5. Run it with sphynx eval ./suite.eval.ts.
An invalid file fails with its path, case and field. A directory is never partially imported.