> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Evaluate an agent

> Create datasets, record evaluation runs, preserve their evidence, and compare outcomes over time.

For a guided start, use [offline evals](/evals/offline). To score completed
sessions, use [online evals](/evals/online). The [Evals overview](/evals/index)
connects both workflows and their cookbooks.

The SDK executes evaluations inside your application or CI process. Agent
Registry stores the durable evidence: the dataset, experiment, immutable
per-case result, scorer definition, trace, and final summary. No Atlan-hosted
runner invokes your model or agent.

```text theme={null}
dataset or inline cases → running experiment → case traces → results → finalize → completed + summary
```

<Frame>
  <img className="block dark:hidden" src="https://mintcdn.com/atlan-602e2b74/h03yFRYof65Y69MX/assets/diagrams/eval-run-light.svg?fit=max&auto=format&n=h03yFRYof65Y69MX&q=85&s=33fe249a103b81076216e904c6c6a887" alt="Evaluation run flow: your application or CI process moves from dataset or inline cases through running experiment, case traces, results, and finalize to completed plus summary. It writes the dataset record, context manifest, experiment result, OTLP trace, versioned scorer, and score span to Agent Registry, which returns summary.scores." width="860" height="718" data-path="assets/diagrams/eval-run-light.svg" />

  <img className="hidden dark:block" src="https://mintcdn.com/atlan-602e2b74/h03yFRYof65Y69MX/assets/diagrams/eval-run-dark.svg?fit=max&auto=format&n=h03yFRYof65Y69MX&q=85&s=41ab3f2106943bf729072457732c18db" alt="Evaluation run flow: your application or CI process moves from dataset or inline cases through running experiment, case traces, results, and finalize to completed plus summary. It writes the dataset record, context manifest, experiment result, OTLP trace, versioned scorer, and score span to Agent Registry, which returns summary.scores." width="860" height="718" data-path="assets/diagrams/eval-run-dark.svg" />
</Frame>

## Run an evaluation

`Eval` is the shortest path. Give it cases, a task, and one or more scorers.
The task and scorers can be synchronous or asynchronous.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import Eval

  run = Eval(
      "Answer quality",
      data=lambda: [
          {"input": "What is 2+2?", "expected": "4"},
          {"input": "What is the capital of France?", "expected": "Paris"},
      ],
      task=call_agent,
      scores=[{
          "name": "accuracy",
          "scorer": lambda output, expected, **_: float(output == expected),
      }],
      config={"model": "candidate-model", "thinking_effort": "medium"},
  )

  print(run.experiment_id)
  ```

  ```typescript TypeScript theme={null}
  import { Eval } from "@atlanai/sdk";

  const run = await Eval("Answer quality", {
    data: () => [
      { input: "What is 2+2?", expected: "4" },
      { input: "What is the capital of France?", expected: "Paris" },
    ],
    task: async (input) => callAgent(input),
    scores: [{
      name: "accuracy",
      scorer: ({ output, expected }) => output === expected ? 1 : 0,
    }],
  }, {
    config: { model: "candidate-model", thinking_effort: "medium" },
  });

  console.log(run.experimentId);
  ```
</CodeGroup>

The default connection uses `ATLAN_API_KEY`, `ATLAN_WORKSPACE_ID`, and
`ATLAN_BASE_URL`. You can pass a configured management client and tracing
logger instead. In an async Python application, call `await async_eval(...)`.

For every case, `Eval`:

1. Opens one root trace and stamps the experiment ID on every child span.
2. Runs the task inside a child span.
3. Resolves an existing Registry scorer by exact generated name, or registers
   one, then records the scorer ID and immutable version on its score span.
4. Flushes the asynchronous trace exporter and reads every root trace ID back
   through the experiment trace filter.
5. Uploads the result with that trace ID and marks the experiment `completed`.
   Registry derives `summary.scores` from the immutable result rows and returns
   it with that same finalization response. A runner, trace verification, or
   upload failure marks the experiment `failed`.

Eval traces always use a sample rate of `1`. If you supply a tracing logger
with a lower sample rate, the run fails before the first case. This prevents a
successful result row from pointing to a sampled-out trace.

Trace read-back waits up to 30 seconds for the trace store by default. Configure
`traceVerificationTimeoutMs` in TypeScript or
`trace_verification_timeout_s` in Python. Disable verification only for an
offline test double; doing so removes the guarantee that every result has a
queryable trace.

<Tip>
  Start with inline `data` while developing. Use one Registry dataset for
  repeatable model comparisons so every experiment runs the same record IDs.
</Tip>

## Start from an existing dataset

Pass a dataset artifact ID or its exact name instead of `data`. The SDK
resolves exactly one dataset, pins its current Registry version, and runs every
record. It fails on a missing or ambiguous name instead of selecting a fuzzy
match.

<CodeGroup>
  ```python Python theme={null}
  run = Eval(
      "Daily agent tasks",
      dataset="daily-agent-tasks",  # exact name, or a dataset_... artifact ID
      task=run_agent,
      scores=[{"name": "task_success", "scorer": score_task_success}],
      config={"model": "candidate-model"},
  )
  ```

  ```typescript TypeScript theme={null}
  const run = await Eval("Daily agent tasks", {
    dataset: "daily-agent-tasks", // exact name, or a dataset_... artifact ID
    task: async (input) => runAgent(input),
    scores: [{ name: "task_success", scorer: scoreTaskSuccess }],
  }, {
    config: { model: "candidate-model" },
  });
  ```
</CodeGroup>

Run each model against that same dataset. The returned experiment ID is the
run identity; each result's `dataset_record_id` is the per-task join across
experiments.

### Dataset versions and readable questions

Set `input_key` when you create a dataset if the human-readable prompt is not
stored under `question`. It defaults to `question`. Registry exposes that value
as `display_input` on every record, so tables and coding agents do not need to
guess which JSON field is the question.

Editing a record creates a new immutable record version. It does not rewrite
the version used by an earlier run. When an experiment starts, Registry stamps
`dataset_snapshot` onto it with the dataset version and every live record's ID,
version, content hash, input, expected value, categories, and provenance. `Eval`
runs from this server-stamped snapshot when available.

Use the record version routes when you need to audit a correction:

```text theme={null}
GET /eval/v1/datasets/{dataset_id}/records/{record_id}/versions
GET /eval/v1/datasets/{dataset_id}/records/{record_id}/versions/{version_ordinal}
```

The second route returns the exact historical question and expected answer.
Nested record and result routes return `404` if the child belongs to a different
dataset or experiment.

### Resume an interrupted upload

Every SDK-created result carries a stable `case_id`. Registry enforces that it
is unique within the experiment. Each create-only result also keeps the input,
expected value, output, per-case scores, and trace ID, so normal result reads
are self-contained even after trace retention. The SDK uses this result
collection as the checkpoint. There is no local progress file to keep in sync.

Give inline cases explicit `id` values when they may be resumed. Dataset-backed
cases automatically use their dataset record IDs.

<CodeGroup>
  ```python Python theme={null}
  run = Eval(
      "Daily agent tasks",
      dataset="daily-agent-tasks",
      task=run_agent,
      scores=scores,
      resume_experiment_id=interrupted_experiment_id,
  )
  ```

  ```typescript TypeScript theme={null}
  const run = await Eval("Daily agent tasks", {
    dataset: "daily-agent-tasks",
    task: runAgent,
    scores,
  }, {
    resumeExperimentId: interruptedExperimentId,
  });
  ```
</CodeGroup>

The experiment must still be `running`, belong to the configured workspace,
use the same dataset, and—when supplied—use the same context-manifest digest.
The SDK lists existing results, skips completed case IDs, uploads only the
remainder, and marks the experiment `completed`. Finalization folds every
immutable result's numeric score snapshot, so a resumed run gets the same
summary without a local progress file or a second API call.

### Summary is finalization output

There is no separate summarize operation. Send runner-observed facts—such as
case counts, elapsed time, or the model actually selected—in the terminal
experiment update. Registry replaces only `summary.scores` with the mean and
count derived from all immutable result rows, preserves the other facts, and
returns the completed experiment with the authoritative summary.

Use the experiment's stored summary for durable run cards and comparisons. Use
experiment trace statistics for live or ad-hoc cost, latency, token, and
post-hoc score-span analysis. Do not aggregate a paginated result page in the
UI; it can produce a partial mean.

### Pin the context that changes behavior

A context manifest is the fingerprint of the agent setup used for a run. It
answers a specific question: which version of every behavior-shaping input was
active when this result was produced?

The manifest contains references and hashes, not the underlying content. The
SDK sorts its entries, writes the canonical manifest to the experiment config,
and computes one manifest digest. Python and TypeScript produce the same digest
for the same entries, regardless of input order.

Each context item has the following shape:

| Field | Required | Example | Rule |
| - | - | - | - |
| `kind` | Yes | `skill` | A lowercase category such as `file`, `prompt`, `skill`, `tool_schema`, `knowledge`, `policy`, `memory`, or `harness`. |
| `name` | Yes | `support-response` | A stable name for this independently versioned input. The pair of `kind` and `name` must be unique in one manifest. |
| `version` | Yes | `git:7d9f2c1` | The pinned revision used by the run. A Git SHA, release version, or snapshot ID works. `latest` is rejected. |
| `digest` | Yes | `sha256:…` | SHA-256 of the exact content or canonical bundle represented by the item. |
| `artifact_id` | No | `skill_01example` | The Registry artifact ID when the input is registered. |
| `version_ordinal` | No | `4` | The exact Registry artifact version, starting at 1. Use it with `artifact_id` when available. |

TypeScript accepts `artifactId` and `versionOrdinal`, then serializes the same
snake-case manifest as Python.

Use one item for each input that can change independently:

| Agent input | Suggested `kind` | What to hash |
| - | - | - |
| Root agent instructions such as `AGENTS.md` or `CLAUDE.md` | `file` | Exact file bytes |
| System or task-routing prompt | `prompt` | Rendered prompt template before task data is inserted |
| Installed skill | `skill` | Canonical skill bundle or Registry artifact version |
| Tool or MCP contract | `tool_schema` | Canonical tool names, descriptions, and input schemas |
| Attached knowledge or retrieval snapshot | `knowledge` | Frozen document set or index snapshot manifest |
| Guardrail or operating policy | `policy` | Exact policy content |
| Shared memory loaded before every case | `memory` | Frozen memory snapshot |
| Agent wrapper or harness configuration | `harness` | Canonical behavior-affecting harness configuration |

Keep run data on its native Eval surface:

| Data | Store it in |
| - | - |
| Task input, expected output, category, and source | Dataset record |
| Model, temperature, thinking effort, and baseline experiment | Experiment config |
| Actual output, error, duration, trace ID, and session ID | Experiment result |
| Scorer definition and output contract | Versioned scorer |
| Score value, explanation, scorer ID, and scorer version | Score span and experiment summary |
| Model calls, tool calls, token usage, and cost | OTLP trace |
| Secrets and credentials | Never the manifest, experiment, result, or trace |

<Note>
  `kind` is extensible, but the item structure is fixed in SDK `0.2.1`.
  Use a new lowercase `kind` for a new artifact category. Do not put arbitrary
  metadata or raw protected content into the manifest.
</Note>

Change an item's `version` and `digest` whenever its effective content changes.
Keep the manifest unchanged when only the dataset case, run timestamp, or model
configuration changes. That separation supports three useful comparisons:

| Hold constant | Change | What the comparison measures |
| - | - | - |
| Dataset and model config | Context manifest | Prompt, skill, tool, or knowledge impact |
| Dataset and context manifest | Model config | Model or reasoning-setting impact |
| Dataset, context manifest, and model config | Nothing material | Run-to-run stability |

<CodeGroup>
  ```python Python theme={null}
  from atlanai import ContextItem, ContextManifest, start_experiment

  context_manifest = ContextManifest([
      ContextItem(
          kind="file",
          name="AGENTS.md",
          version=release_commit,
          digest=agent_instructions_digest,
      ),
      ContextItem(
          kind="skill",
          name="support-response",
          version="4",
          digest=skill_digest,
          artifact_id="skill_01example",
          version_ordinal=4,
      ),
      ContextItem(
          kind="tool_schema",
          name="support-tools",
          version=tool_contract_commit,
          digest=tool_schema_digest,
      ),
  ])

  run = start_experiment(
      client,
      "daily-agent-tasks",  # exact name, or a dataset_... artifact ID
      {"name": "candidate-run", "config": {"model": "candidate-model"}},
      context_manifest=context_manifest,
  )

  with run.trace():
      output = existing_runner()

  print(run.experiment_id)
  ```

  ```typescript TypeScript theme={null}
  import {
    createContextManifest,
    startExperiment,
  } from "@atlanai/sdk";
  import { propagateAttributes } from "@atlanai/tools/sdk/how-tos/tracing";

  const contextManifest = await createContextManifest([{
    kind: "file",
    name: "AGENTS.md",
    version: releaseCommit,
    digest: agentInstructionsDigest,
  }, {
    kind: "skill",
    name: "support-response",
    version: "4",
    digest: skillDigest,
    artifactId: "skill_01example",
    versionOrdinal: 4,
  }, {
    kind: "tool_schema",
    name: "support-tools",
    version: toolContractCommit,
    digest: toolSchemaDigest,
  }]);

  const run = await startExperiment(
    client,
    "daily-agent-tasks", // exact name, or a dataset_... artifact ID
    { name: "candidate-run", config: { model: "candidate-model" } },
    { contextManifest },
  );

  const output = await propagateAttributes(
    run.traceOptions,
    () => existingRunner(),
  );

  console.log(run.experimentId);
  ```
</CodeGroup>

`run.experiment` is the generated experiment-create response. Python exposes
its ID as `run.id` and `run.experiment_id`; TypeScript exposes `run.id` and
`run.experimentId`. A name match is exact and workspace-scoped. Multiple exact
matches fail instead of selecting one.

The experiment config keeps both `context_manifest` and
`context_manifest_digest`. The trace scope carries the digest alongside the
experiment ID. This gives every result a path back to the exact prompt, skill,
tool contract, and knowledge versions that shaped it.

<Note>
  `start_experiment` only starts the Registry lifecycle for an existing
  harness. Use `Eval` when the SDK should execute cases, upload results,
  and finalize the experiment for you.
</Note>

## Control the lifecycle directly

Start by creating a dataset and the cases it contains. Every Eval create body
needs `workspace_id`, even when the client has a default workspace header.

```python Python theme={null}
from atlanai import AtlanClient

workspace_id = "workspace_01example"
client = AtlanClient(
    "https://gateway.example",
    bearer_token="...",
    workspace=workspace_id,
)

dataset = client.datasets.create({
    "workspace_id": workspace_id,
    "name": "support-quality-set",
    "display_name": "Support quality set",
    "extra": {"data_snapshot_ref": "support-sample-2026-09"},
})

record = client.datasets.records.create(dataset.id, {
    "workspace_id": workspace_id,
    "name": "reset-password-case",
    "input": {"question": "How do I reset my password?"},
    "expected": {"must_include": ["reset link", "security guidance"]},
    "source_kind": "manual",
    "categories": ["support", "account"],
})

scorer = client.scorers.create({
    "workspace_id": workspace_id,
    "name": "support-answer-quality",
    "scorer_kind": "code",
    "scope": "result",
    "spec": {"entrypoint": "evaluate_support_answer"},
    "outputs": {"quality": {"type": "numeric", "min": 0, "max": 1}},
})

experiment = client.experiments.create({
    "workspace_id": workspace_id,
    "name": "support-agent-candidate",
    "dataset_id": dataset.id,
    "subject_kind": "agent",
    "subject_id": "agent_01example",
    "config": {"model": "candidate-model", "prompt_version": "v2"},
})

# The external runner executes the case and emits its trace and score spans.
result = client.experiments.results.create(experiment.id, {
    "workspace_id": workspace_id,
    "name": "reset-password-result",
    "dataset_record_id": record.id,
    "input": {"question": "How do I reset my password?"},
    "expected": {"must_include": ["reset link", "security guidance"]},
    "output": {"answer": "Use the reset-password link on the sign-in page."},
    "scores": {"quality": 0.9},
    "duration_ms": 410,
    "trace_id": trace_id_from_runner,
})

# One write seals the run and returns Registry's durable score summary.
experiment = client.experiments.update(experiment.id, {
    "experiment_status": "completed",
    "summary": {
        "metrics": {"cases": 1, "duration_ms": 410},
        "effective_model": "candidate-model",
    },
})
print(experiment.summary["scores"])
```

The TypeScript client has the same resource tree, with `camelCase` body keys
and awaited calls: `await client.datasets.records.create(datasetId, body)` and
`await client.experiments.results.create(experimentId, body)`.

An experiment's trace views scope OTel data by the
`atlan.eval.experiment_id` span attribute. `run.trace()` in Python and
`propagateAttributes(run.traceOptions, fn)` in TypeScript stamp that association
and the context-manifest digest on every span created inside the runner scope.

## Keep evidence that remains useful

The value of an evaluation is being able to diagnose a regression months later,
not only its average score. Preserve these fields as part of each run:

| Keep | Why it matters later | Eval surface |
| - | - | - |
| Input, expected output, categories, and source reference | Reproduce the case and slice failures by scenario or provenance. | Dataset record |
| Dataset snapshot reference | Know which frozen source data a benchmark actually used. | Dataset `extra.data_snapshot_ref` |
| Subject, model/prompt/configuration, and baseline experiment | Compare candidates fairly and explain why a score moved. | Experiment |
| Actual output, duration, error, trace ID, and session ID | Move from a failed row to the exact execution and its latency or failure. | Experiment result |
| Scorer kind, scope, definition, and output schema | Keep the interpretation of a score stable as evaluators evolve. | Scorer |
| Trace-level model calls, tool calls, token usage, cost, and score comments | Diagnose whether a quality shift came from the model, tools, prompt, cost, or the scorer. | Tracing SDK and experiment trace views |

Scorers are versioned artifacts. Results are create-only, and a completed or
failed experiment is frozen. That combination preserves the run as evidence
instead of allowing later dataset edits or result updates to rewrite history.

## Read and compare

Use the resource tree to review an experiment's durable results and the live
trace detail behind them:

```python Python theme={null}
results = client.experiments.results.list(experiment.id)
traces = client.experiments.traces.list(experiment.id)
one_trace = client.experiments.traces.get(experiment.id, trace_id_from_runner)
spans = client.experiments.traces.list_spans(experiment.id, trace_id_from_runner)
```

For the complete endpoint list, including search, bulk result upload, archive,
and trace statistics, see the generated resource references below.

<CardGroup cols={2}>
  <Card title="Datasets" icon="list" href="/tools/sdk/references/operations/datasets">
    Curate dataset records and preserve their provenance.
  </Card>

  <Card title="Experiments" icon="flask" href="/tools/sdk/references/operations/experiments">
    Record runs, results, traces, and score rollups.
  </Card>

  <Card title="Scorers" icon="star" href="/tools/sdk/references/operations/scorers">
    Version the score definition alongside the evaluation.
  </Card>

  <Card title="Tracing and scores" icon="diagram-project" href="/tools/sdk/how-tos/tracing">
    Capture the execution evidence that makes an evaluation explainable.
  </Card>
</CardGroup>
