Eval runs your task against each case, calls the score functions you provide, exports a trace for each case, and stores the results in an experiment. The gateway stores the evidence; your process executes the task.
Set up
Install the SDK with tracing. Python needs the tracing extra; TypeScript includes tracing in the package.Run two cases
This example deliberately gets one answer wrong so you can see a failed case. Replaceanswer with your agent call once the flow works.
summary.scores values are derived from stored case results when the experiment completes. Inspect a low-scoring case’s trace before changing the task or its scorer. The result tells you which case failed; its trace helps explain why.
Reuse a dataset
Inline cases are useful while developing. Put stable cases in a Registry dataset when you need to compare runs. Each record’sinput is a JSON object. By default, input.question becomes the value passed to task; a single-key expected: { value: ... } becomes the scalar passed to the scorer.
The SDK’s push_dataset and pushDataset helpers create a dataset once, add new records, and update changed records without creating an empty new version on every push.
config or a context manifest.
Pin the context that changed
A context manifest records immutable versions and SHA-256 digests for behavior-shaping inputs such as instructions, prompts, skills, tool schemas, or retrieval snapshots. The SDK writes the manifest and its digest into the experiment config. The gateway stores them but does not resolve or validate the referenced content, so point each entry at a version you can still retrieve.latest is not a reproducible version.
Resume a stopped run
Save the experiment ID as soon as the run starts withon_start in Python or onStart in TypeScript. Restart with the same dataset, task, scorers, configuration, and context manifest, plus resume_experiment_id or resumeExperimentId. The SDK skips case IDs already stored and finishes the remaining cases.
For example, the first run can write the ID with on_start=lambda event: Path("eval-run-id.txt").write_text(event["experiment_id"]) in Python, or onStart: ({ experimentId }) => writeFileSync("eval-run-id.txt", experimentId) in TypeScript. Import Path from pathlib or writeFileSync from node:fs, respectively. In a later process, read that ID and resume:
running so it can be resumed. A task or scorer failure is recorded on that case and the runner continues. A remote task with side effects needs its own idempotency or reconciliation before retrying it.