Run an evaluation
Eval is the shortest path. Give it cases, a task, and one or more scorers.
The task and scorers can be synchronous or asynchronous.
ATLAN_API_KEY, ATLAN_WORKSPACE_ID, and
ATLAN_BASE_URL. You can pass a configured management client and tracing
logger instead. In an async Python application, call await async_eval(...).
For every case, Eval:
- Opens one root trace and stamps the experiment ID on every child span.
- Runs the task inside a child span.
- Resolves an existing Registry scorer by exact generated name, or registers one, then records the scorer ID and immutable version on its score span.
- Flushes the asynchronous trace exporter and reads every root trace ID back through the experiment trace filter.
- Uploads the result with that trace ID and marks the experiment
completed. Registry derivessummary.scoresfrom the immutable result rows and returns it with that same finalization response. A runner, trace verification, or upload failure marks the experimentfailed.
1. If you supply a tracing logger
with a lower sample rate, the run fails before the first case. This prevents a
successful result row from pointing to a sampled-out trace.
Trace read-back waits up to 30 seconds for the trace store by default. Configure
traceVerificationTimeoutMs in TypeScript or
trace_verification_timeout_s in Python. Disable verification only for an
offline test double; doing so removes the guarantee that every result has a
queryable trace.
Start from an existing dataset
Pass a dataset artifact ID or its exact name instead ofdata. The SDK
resolves exactly one dataset, pins its current Registry version, and runs every
record. It fails on a missing or ambiguous name instead of selecting a fuzzy
match.
dataset_record_id is the per-task join across
experiments.
Dataset versions and readable questions
Setinput_key when you create a dataset if the human-readable prompt is not
stored under question. It defaults to question. Registry exposes that value
as display_input on every record, so tables and coding agents do not need to
guess which JSON field is the question.
Editing a record creates a new immutable record version. It does not rewrite
the version used by an earlier run. When an experiment starts, Registry stamps
dataset_snapshot onto it with the dataset version and every live record’s ID,
version, content hash, input, expected value, categories, and provenance. Eval
runs from this server-stamped snapshot when available.
Use the record version routes when you need to audit a correction:
404 if the child belongs to a different
dataset or experiment.
Resume an interrupted upload
Every SDK-created result carries a stablecase_id. Registry enforces that it
is unique within the experiment. Each create-only result also keeps the input,
expected value, output, per-case scores, and trace ID, so normal result reads
are self-contained even after trace retention. The SDK uses this result
collection as the checkpoint. There is no local progress file to keep in sync.
Give inline cases explicit id values when they may be resumed. Dataset-backed
cases automatically use their dataset record IDs.
running, belong to the configured workspace,
use the same dataset, and—when supplied—use the same context-manifest digest.
The SDK lists existing results, skips completed case IDs, uploads only the
remainder, and marks the experiment completed. Finalization folds every
immutable result’s numeric score snapshot, so a resumed run gets the same
summary without a local progress file or a second API call.
Summary is finalization output
There is no separate summarize operation. Send runner-observed facts—such as case counts, elapsed time, or the model actually selected—in the terminal experiment update. Registry replaces onlysummary.scores with the mean and
count derived from all immutable result rows, preserves the other facts, and
returns the completed experiment with the authoritative summary.
Use the experiment’s stored summary for durable run cards and comparisons. Use
experiment trace statistics for live or ad-hoc cost, latency, token, and
post-hoc score-span analysis. Do not aggregate a paginated result page in the
UI; it can produce a partial mean.
Pin the context that changes behavior
A context manifest is the fingerprint of the agent setup used for a run. It answers a specific question: which version of every behavior-shaping input was active when this result was produced? The manifest contains references and hashes, not the underlying content. The SDK sorts its entries, writes the canonical manifest to the experiment config, and computes one manifest digest. Python and TypeScript produce the same digest for the same entries, regardless of input order. Each context item has the following shape:
TypeScript accepts
artifactId and versionOrdinal, then serializes the same
snake-case manifest as Python.
Use one item for each input that can change independently:
Keep run data on its native Eval surface:
kind is extensible, but the item structure is fixed in SDK 0.2.1.
Use a new lowercase kind for a new artifact category. Do not put arbitrary
metadata or raw protected content into the manifest.version and digest whenever its effective content changes.
Keep the manifest unchanged when only the dataset case, run timestamp, or model
configuration changes. That separation supports three useful comparisons:
run.experiment is the generated experiment-create response. Python exposes
its ID as run.id and run.experiment_id; TypeScript exposes run.id and
run.experimentId. A name match is exact and workspace-scoped. Multiple exact
matches fail instead of selecting one.
The experiment config keeps both context_manifest and
context_manifest_digest. The trace scope carries the digest alongside the
experiment ID. This gives every result a path back to the exact prompt, skill,
tool contract, and knowledge versions that shaped it.
start_experiment only starts the Registry lifecycle for an existing
harness. Use Eval when the SDK should execute cases, upload results,
and finalize the experiment for you.Control the lifecycle directly
Start by creating a dataset and the cases it contains. Every Eval create body needsworkspace_id, even when the client has a default workspace header.
Python
camelCase body keys
and awaited calls: await client.datasets.records.create(datasetId, body) and
await client.experiments.results.create(experimentId, body).
An experiment’s trace views scope OTel data by the
atlan.eval.experiment_id span attribute. run.trace() in Python and
propagateAttributes(run.traceOptions, fn) in TypeScript stamp that association
and the context-manifest digest on every span created inside the runner scope.
Keep evidence that remains useful
The value of an evaluation is being able to diagnose a regression months later, not only its average score. Preserve these fields as part of each run:
Scorers are versioned artifacts. Results are create-only, and a completed or
failed experiment is frozen. That combination preserves the run as evidence
instead of allowing later dataset edits or result updates to rewrite history.
Read and compare
Use the resource tree to review an experiment’s durable results and the live trace detail behind them:Python
Datasets
Curate dataset records and preserve their provenance.
Experiments
Record runs, results, traces, and score rollups.
Scorers
Version the score definition alongside the evaluation.
Tracing and scores
Capture the execution evidence that makes an evaluation explainable.