Skip to main content
Run offline evals before changing an agent or model in production. Eval runs your task against each case, calls the score functions you provide, exports a trace for each case, and stores the results in an experiment. The gateway stores the evidence; your process executes the task.

Set up

Install the SDK with tracing. Python needs the tracing extra; TypeScript includes tracing in the package.
Provide a workspace-scoped credential through the process environment:
The credential must be able to create datasets, scorers, experiments, and results in the workspace. Keep it out of source files and logs.

Run two cases

This example deliberately gets one answer wrong so you can see a failed case. Replace answer with your agent call once the flow works.
The summary.scores values are derived from stored case results when the experiment completes. Inspect a low-scoring case’s trace before changing the task or its scorer. The result tells you which case failed; its trace helps explain why.

Reuse a dataset

Inline cases are useful while developing. Put stable cases in a Registry dataset when you need to compare runs. Each record’s input is a JSON object. By default, input.question becomes the value passed to task; a single-key expected: { value: ... } becomes the scalar passed to the scorer. The SDK’s push_dataset and pushDataset helpers create a dataset once, add new records, and update changed records without creating an empty new version on every push.
Use the same dataset ID for the baseline and candidate. Each experiment pins a dataset snapshot at creation, so later record edits do not change an earlier run. Record model, prompt, agent, and tool revisions in the experiment config or a context manifest.

Pin the context that changed

A context manifest records immutable versions and SHA-256 digests for behavior-shaping inputs such as instructions, prompts, skills, tool schemas, or retrieval snapshots. The SDK writes the manifest and its digest into the experiment config. The gateway stores them but does not resolve or validate the referenced content, so point each entry at a version you can still retrieve.
Replace the example digest with the digest of the actual content. A label such as latest is not a reproducible version.

Resume a stopped run

Save the experiment ID as soon as the run starts with on_start in Python or onStart in TypeScript. Restart with the same dataset, task, scorers, configuration, and context manifest, plus resume_experiment_id or resumeExperimentId. The SDK skips case IDs already stored and finishes the remaining cases. For example, the first run can write the ID with on_start=lambda event: Path("eval-run-id.txt").write_text(event["experiment_id"]) in Python, or onStart: ({ experimentId }) => writeFileSync("eval-run-id.txt", experimentId) in TypeScript. Import Path from pathlib or writeFileSync from node:fs, respectively. In a later process, read that ID and resume:
Run only one writer per experiment. A failed trace export, verification, or result upload leaves the experiment running so it can be resumed. A task or scorer failure is recorded on that case and the runner continues. A remote task with side effects needs its own idempotency or reconciliation before retrying it.

Read the result

The experiment response contains the durable score summary. Read its stored results and traces for case-level diagnosis:
For a model comparison with a shared dataset, use the compare versions cookbook. For framework instrumentation and manual score spans, see tracing integrations and scores.