Skip to main content
For a guided start, use offline evals. To score completed sessions, use online evals. The Evals overview connects both workflows and their cookbooks. The SDK executes evaluations inside your application or CI process. Agent Registry stores the durable evidence: the dataset, experiment, immutable per-case result, scorer definition, trace, and final summary. No Atlan-hosted runner invokes your model or agent.
Evaluation run flow: your application or CI process moves from dataset or inline cases through running experiment, case traces, results, and finalize to completed plus summary. It writes the dataset record, context manifest, experiment result, OTLP trace, versioned scorer, and score span to Agent Registry, which returns summary.scores.Evaluation run flow: your application or CI process moves from dataset or inline cases through running experiment, case traces, results, and finalize to completed plus summary. It writes the dataset record, context manifest, experiment result, OTLP trace, versioned scorer, and score span to Agent Registry, which returns summary.scores.

Run an evaluation

Eval is the shortest path. Give it cases, a task, and one or more scorers. The task and scorers can be synchronous or asynchronous.
The default connection uses ATLAN_API_KEY, ATLAN_WORKSPACE_ID, and ATLAN_BASE_URL. You can pass a configured management client and tracing logger instead. In an async Python application, call await async_eval(...). For every case, Eval:
  1. Opens one root trace and stamps the experiment ID on every child span.
  2. Runs the task inside a child span.
  3. Resolves an existing Registry scorer by exact generated name, or registers one, then records the scorer ID and immutable version on its score span.
  4. Flushes the asynchronous trace exporter and reads every root trace ID back through the experiment trace filter.
  5. Uploads the result with that trace ID and marks the experiment completed. Registry derives summary.scores from the immutable result rows and returns it with that same finalization response. A runner, trace verification, or upload failure marks the experiment failed.
Eval traces always use a sample rate of 1. If you supply a tracing logger with a lower sample rate, the run fails before the first case. This prevents a successful result row from pointing to a sampled-out trace. Trace read-back waits up to 30 seconds for the trace store by default. Configure traceVerificationTimeoutMs in TypeScript or trace_verification_timeout_s in Python. Disable verification only for an offline test double; doing so removes the guarantee that every result has a queryable trace.
Start with inline data while developing. Use one Registry dataset for repeatable model comparisons so every experiment runs the same record IDs.

Start from an existing dataset

Pass a dataset artifact ID or its exact name instead of data. The SDK resolves exactly one dataset, pins its current Registry version, and runs every record. It fails on a missing or ambiguous name instead of selecting a fuzzy match.
Run each model against that same dataset. The returned experiment ID is the run identity; each result’s dataset_record_id is the per-task join across experiments.

Dataset versions and readable questions

Set input_key when you create a dataset if the human-readable prompt is not stored under question. It defaults to question. Registry exposes that value as display_input on every record, so tables and coding agents do not need to guess which JSON field is the question. Editing a record creates a new immutable record version. It does not rewrite the version used by an earlier run. When an experiment starts, Registry stamps dataset_snapshot onto it with the dataset version and every live record’s ID, version, content hash, input, expected value, categories, and provenance. Eval runs from this server-stamped snapshot when available. Use the record version routes when you need to audit a correction:
The second route returns the exact historical question and expected answer. Nested record and result routes return 404 if the child belongs to a different dataset or experiment.

Resume an interrupted upload

Every SDK-created result carries a stable case_id. Registry enforces that it is unique within the experiment. Each create-only result also keeps the input, expected value, output, per-case scores, and trace ID, so normal result reads are self-contained even after trace retention. The SDK uses this result collection as the checkpoint. There is no local progress file to keep in sync. Give inline cases explicit id values when they may be resumed. Dataset-backed cases automatically use their dataset record IDs.
The experiment must still be running, belong to the configured workspace, use the same dataset, and—when supplied—use the same context-manifest digest. The SDK lists existing results, skips completed case IDs, uploads only the remainder, and marks the experiment completed. Finalization folds every immutable result’s numeric score snapshot, so a resumed run gets the same summary without a local progress file or a second API call.

Summary is finalization output

There is no separate summarize operation. Send runner-observed facts—such as case counts, elapsed time, or the model actually selected—in the terminal experiment update. Registry replaces only summary.scores with the mean and count derived from all immutable result rows, preserves the other facts, and returns the completed experiment with the authoritative summary. Use the experiment’s stored summary for durable run cards and comparisons. Use experiment trace statistics for live or ad-hoc cost, latency, token, and post-hoc score-span analysis. Do not aggregate a paginated result page in the UI; it can produce a partial mean.

Pin the context that changes behavior

A context manifest is the fingerprint of the agent setup used for a run. It answers a specific question: which version of every behavior-shaping input was active when this result was produced? The manifest contains references and hashes, not the underlying content. The SDK sorts its entries, writes the canonical manifest to the experiment config, and computes one manifest digest. Python and TypeScript produce the same digest for the same entries, regardless of input order. Each context item has the following shape: TypeScript accepts artifactId and versionOrdinal, then serializes the same snake-case manifest as Python. Use one item for each input that can change independently: Keep run data on its native Eval surface:
kind is extensible, but the item structure is fixed in SDK 0.2.1. Use a new lowercase kind for a new artifact category. Do not put arbitrary metadata or raw protected content into the manifest.
Change an item’s version and digest whenever its effective content changes. Keep the manifest unchanged when only the dataset case, run timestamp, or model configuration changes. That separation supports three useful comparisons:
run.experiment is the generated experiment-create response. Python exposes its ID as run.id and run.experiment_id; TypeScript exposes run.id and run.experimentId. A name match is exact and workspace-scoped. Multiple exact matches fail instead of selecting one. The experiment config keeps both context_manifest and context_manifest_digest. The trace scope carries the digest alongside the experiment ID. This gives every result a path back to the exact prompt, skill, tool contract, and knowledge versions that shaped it.
start_experiment only starts the Registry lifecycle for an existing harness. Use Eval when the SDK should execute cases, upload results, and finalize the experiment for you.

Control the lifecycle directly

Start by creating a dataset and the cases it contains. Every Eval create body needs workspace_id, even when the client has a default workspace header.
Python
The TypeScript client has the same resource tree, with camelCase body keys and awaited calls: await client.datasets.records.create(datasetId, body) and await client.experiments.results.create(experimentId, body). An experiment’s trace views scope OTel data by the atlan.eval.experiment_id span attribute. run.trace() in Python and propagateAttributes(run.traceOptions, fn) in TypeScript stamp that association and the context-manifest digest on every span created inside the runner scope.

Keep evidence that remains useful

The value of an evaluation is being able to diagnose a regression months later, not only its average score. Preserve these fields as part of each run: Scorers are versioned artifacts. Results are create-only, and a completed or failed experiment is frozen. That combination preserves the run as evidence instead of allowing later dataset edits or result updates to rewrite history.

Read and compare

Use the resource tree to review an experiment’s durable results and the live trace detail behind them:
Python
For the complete endpoint list, including search, bulk result upload, archive, and trace statistics, see the generated resource references below.

Datasets

Curate dataset records and preserve their provenance.

Experiments

Record runs, results, traces, and score rollups.

Scorers

Version the score definition alongside the evaluation.

Tracing and scores

Capture the execution evidence that makes an evaluation explainable.