> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Run an offline eval

> Run a fixed set of cases with the Python or TypeScript SDK and inspect trace-linked results.

Run offline evals before changing an agent or model in production. `Eval` runs your task against each case, calls the score functions you provide, exports a trace for each case, and stores the results in an experiment. The gateway stores the evidence; your process executes the task.

## Set up

Install the SDK with tracing. Python needs the tracing extra; TypeScript includes tracing in the package.

<CodeGroup>
  ```bash Python theme={null}
  pip install 'atlanai[tracing]'
  ```

  ```bash TypeScript theme={null}
  npm install @atlanai/sdk
  ```
</CodeGroup>

Provide a workspace-scoped credential through the process environment:

```bash theme={null}
export ATLAN_API_KEY="<token-from-your-secret-store>"
export ATLAN_WORKSPACE_ID="workspace_01example"
# Optional when using a non-default Gateway:
export ATLAN_BASE_URL="https://<gateway-host>"
```

The credential must be able to create datasets, scorers, experiments, and results in the workspace. Keep it out of source files and logs.

## Run two cases

This example deliberately gets one answer wrong so you can see a failed case. Replace `answer` with your agent call once the flow works.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import Eval


  def answer(question: str) -> str:
      return "4"  # Replace with your agent call.


  def accuracy(*, output: str, expected: str, **_) -> float:
      return float(output == expected)


  run = Eval(
      "answer-regression",
      data=lambda: [
          {"id": "arithmetic", "input": "What is 2 + 2?", "expected": "4"},
          {"id": "geography", "input": "What is the capital of France?", "expected": "Paris"},
      ],
      task=answer,
      scores=[accuracy],
      config={"candidate": "example-v1"},
  )

  print(run.experiment_id, run.summary)
  for case in run.results:
      print(case.case_id, case.scores, case.trace_id)
  ```

  ```typescript TypeScript theme={null}
  import { Eval } from "@atlanai/sdk";

  const run = await Eval("answer-regression", {
    data: () => [
      { id: "arithmetic", input: "What is 2 + 2?", expected: "4" },
      { id: "geography", input: "What is the capital of France?", expected: "Paris" },
    ],
    task: async (_question) => "4", // Replace with your agent call.
    scores: [{
      name: "accuracy",
      scorer: ({ output, expected }) => Number(output === expected),
    }],
  }, { config: { candidate: "example-v1" } });

  console.log(run.experimentId, run.summary);
  for (const caseResult of run.results) {
    console.log(caseResult.caseId, caseResult.scores, caseResult.traceId);
  }
  ```
</CodeGroup>

The `summary.scores` values are derived from stored case results when the experiment completes. Inspect a low-scoring case's trace before changing the task or its scorer. The result tells you *which* case failed; its trace helps explain *why*.

## Reuse a dataset

Inline cases are useful while developing. Put stable cases in a Registry dataset when you need to compare runs. Each record's `input` is a JSON object. By default, `input.question` becomes the value passed to `task`; a single-key `expected: { value: ... }` becomes the scalar passed to the scorer.

The SDK's `push_dataset` and `pushDataset` helpers create a dataset once, add new records, and update changed records without creating an empty new version on every push.

<CodeGroup>
  ```python Python theme={null}
  import os
  from atlanai import AtlanClient, Eval, push_dataset

  client = AtlanClient(
      os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
      bearer_token=os.environ["ATLAN_API_KEY"],
      workspace=os.environ["ATLAN_WORKSPACE_ID"],
  )

  suite = push_dataset(client, "answer-regression-cases", [
      {"name": "arithmetic", "input": {"question": "What is 2 + 2?"}, "expected": {"value": "4"}},
      {"name": "geography", "input": {"question": "What is the capital of France?"}, "expected": {"value": "Paris"}},
  ])

  run = Eval("candidate-v2", dataset=suite.dataset.id, task=answer, scores=[accuracy])
  print(run.experiment_id, run.summary)
  ```

  ```typescript TypeScript theme={null}
  import { AtlanClient, Eval, pushDataset } from "@atlanai/sdk";

  const client = new AtlanClient({
    gatewayOrigin: process.env.ATLAN_BASE_URL ?? "https://api.atlan.com",
    bearerToken: process.env.ATLAN_API_KEY!,
    workspace: process.env.ATLAN_WORKSPACE_ID!,
  });

  const suite = await pushDataset(client, "answer-regression-cases", [
    { name: "arithmetic", input: { question: "What is 2 + 2?" }, expected: { value: "4" } },
    { name: "geography", input: { question: "What is the capital of France?" }, expected: { value: "Paris" } },
  ]);

  const evaluator = {
    dataset: suite.id,
    task: async (_question: string) => "4",
    scores: [{
      name: "accuracy",
      scorer: ({ output, expected }: { output: string; expected: string }) => Number(output === expected),
    }],
  };
  const run = await Eval("candidate-v2", evaluator, { config: { candidate: "example-v2" } });
  console.log(run.experimentId, run.summary);
  ```
</CodeGroup>

Use the same dataset ID for the baseline and candidate. Each experiment pins a dataset snapshot at creation, so later record edits do not change an earlier run. Record model, prompt, agent, and tool revisions in the experiment `config` or a context manifest.

### Pin the context that changed

A context manifest records immutable versions and SHA-256 digests for behavior-shaping inputs such as instructions, prompts, skills, tool schemas, or retrieval snapshots. The SDK writes the manifest and its digest into the experiment config. The gateway stores them but does not resolve or validate the referenced content, so point each entry at a version you can still retrieve.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import ContextItem, ContextManifest

  context = ContextManifest([ContextItem(
      kind="prompt", name="answer-system", version="git:0123456789abcdef0123456789abcdef01234567",
      digest="sha256:" + "0" * 64,
  )])
  run = Eval("candidate-v2", dataset=suite.dataset.id, task=answer,
             scores=[accuracy], context_manifest=context)
  ```

  ```typescript TypeScript theme={null}
  import { createContextManifest } from "@atlanai/sdk";

  const contextManifest = await createContextManifest([{
    kind: "prompt", name: "answer-system",
    version: "git:0123456789abcdef0123456789abcdef01234567",
    digest: `sha256:${"0".repeat(64)}`,
  }]);
  const manifestRun = await Eval("candidate-v2", evaluator, {
    config: { candidate: "example-v2" },
    contextManifest,
  });
  ```
</CodeGroup>

Replace the example digest with the digest of the actual content. A label such as `latest` is not a reproducible version.

## Resume a stopped run

Save the experiment ID as soon as the run starts with `on_start` in Python or `onStart` in TypeScript. Restart with the same dataset, task, scorers, configuration, and context manifest, plus `resume_experiment_id` or `resumeExperimentId`. The SDK skips case IDs already stored and finishes the remaining cases.

For example, the first run can write the ID with `on_start=lambda event: Path("eval-run-id.txt").write_text(event["experiment_id"])` in Python, or `onStart: ({ experimentId }) => writeFileSync("eval-run-id.txt", experimentId)` in TypeScript. Import `Path` from `pathlib` or `writeFileSync` from `node:fs`, respectively. In a later process, read that ID and resume:

<CodeGroup>
  ```python Python theme={null}
  from pathlib import Path

  saved_experiment_id = Path("eval-run-id.txt").read_text().strip()
  resumed = Eval(
      "candidate-v2", dataset=suite.dataset.id, task=answer, scores=[accuracy],
      config={"candidate": "example-v2"},
      resume_experiment_id=saved_experiment_id,
  )
  ```

  ```typescript TypeScript theme={null}
  import { readFileSync } from "node:fs";

  const savedExperimentId = readFileSync("eval-run-id.txt", "utf8").trim();
  const resumed = await Eval("candidate-v2", evaluator, {
    config: { candidate: "example-v2" },
    resumeExperimentId: savedExperimentId,
  });
  ```
</CodeGroup>

Run only one writer per experiment. A failed trace export, verification, or result upload leaves the experiment `running` so it can be resumed. A task or scorer failure is recorded on that case and the runner continues. A remote task with side effects needs its own idempotency or reconciliation before retrying it.

## Read the result

The experiment response contains the durable score summary. Read its stored results and traces for case-level diagnosis:

<CodeGroup>
  ```python Python theme={null}
  experiment = client.experiments.get(run.experiment_id)
  results = client.experiments.results.list(run.experiment_id)
  traces = client.experiments.traces.list(run.experiment_id)
  print(experiment.summary, len(results.items), len(traces.items))
  ```

  ```typescript TypeScript theme={null}
  const experiment = await client.experiments.get(run.experimentId);
  const results = await client.experiments.results.list(run.experimentId);
  const traces = await client.experiments.traces.list(run.experimentId);
  console.log(experiment.summary, results.items.length, traces.items.length);
  ```
</CodeGroup>

For a model comparison with a shared dataset, use the [compare versions cookbook](/evals/cookbooks/compare-versions). For framework instrumentation and manual score spans, see [tracing integrations](/tools/sdk/how-tos/integrations) and [scores](/tools/sdk/how-tos/scores).
