> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Concepts

> Datasets, experiments, results, scorers, score spans, and insights: what each object holds, what is versioned, and what is frozen.

Every eval object answers one question. Knowing which objects are versioned, create-only, or frozen tells you what you can rely on when you compare runs months apart.

```text theme={null}
dataset ──has──► dataset records (versioned)
   │
   └─ pinned by ─► experiment ──has──► experiment results (create-only)
                      │                     │
                      │                     └─ trace_id ─► trace ──has──► score spans
                      │                                                    │
                      └─ summary.scores ◄── folded from results            └─ cite ─► scorer version
```

## The objects

| Object | What it is | Mutability |
| - | - | - |
| **Dataset** | A named collection of cases in a workspace. | Editable. Its records are versioned. |
| **Dataset record** | One case: `input` (a JSON object), optional `expected`, `categories`, a `label`, `extra` metadata, and provenance (`source_kind`, `source_ref`). | Editing a record appends a new immutable version. |
| **Experiment** | One run of a task over a dataset or inline cases. It holds the run `config`, an optional `baseline_experiment_id`, an optional subject (the agent, harness, or skill under test), a status, and a `summary`. | `running` accepts results. Once `completed` or `failed` it is frozen, and any change is refused with `409`. |
| **Experiment result** | One case's outcome: `case_id`, `input`, `expected`, `output`, `scores`, `error`, `duration_ms`, `trace_id`, and the `dataset_record_id` it ran. | Create-only. It can never be edited or deleted. |
| **Scorer** | The definition of what a score means. `code` scorers are functions your runner executes; `llm_judge` scorers are hosted judge definitions Atlan executes. `human` scorers catalog a manual review. | Every edit through the eval API appends a version. Scores cite the exact version. |
| **Trace** | The spans your task produced for one case: model calls, tool calls, tokens, cost, latency, errors. | Written by the tracing SDK and kept for your trace retention period. |
| **Score span** | A child span named `score.<name>` carrying the score value and the scorer ID and version that produced it. | Written once, with the trace. |
| **Session insight** | One hosted judge's answer about one production session or trace. | Re-scoring the same session with the same scorer rewrites it. |

## Snapshots make runs comparable

When an experiment starts from a dataset, Atlan stamps a **`dataset_snapshot`** onto it: the dataset version plus each live record's ID, version, content hash, input, and expected value. The run executes from that snapshot. Editing a record later changes only *future* experiments, so two runs on the same dataset compare the same cases unless you deliberately changed them.

Results join across experiments on **`dataset_record_id`**. That is the key for "which cases got better or worse".

## Where the summary comes from

`summary.scores` is computed by Atlan, not by your runner. When the experiment moves from `running` to `completed` or `failed`, Atlan reads every stored result and replaces `summary.scores` with `{name: {mean, count}}` for each score name. A `null` score, meaning "not applicable", is left out of both the mean and the count. Other summary fields your runner sends, such as case counts and duration, are kept as sent.

Use the stored summary for release decisions and run history. Do not average a page of results yourself: a partial page gives a partial mean.

## Offline and online scores are stored differently

| | Offline (`Eval`) | Online (hosted judge) |
| - | - | - |
| Score value | On the result row (`scores`) and as a score span on the case trace | On a session insight |
| Scorer kind | `code`, run in your process | `llm_judge`, run by Atlan |
| Aggregate | `summary.scores` on the experiment | Per-answer rows; compare them in the Atlan app |

A score span you write from application code, for example a user's thumbs-up, is a third path. See [scores from your application](/evals/production-scores).

## What Atlan does and does not do

* **Atlan stores and computes.** It stores datasets, snapshots, results, scorer versions, and traces. It computes summaries. It executes hosted `llm_judge` scorers over recorded sessions.
* **Atlan does not run your agent.** Offline evals execute in your process or CI job. There is no hosted runner, and no Atlan service calls your model or tools.
* **There is no compare endpoint.** A baseline is a stored pointer. You compare two experiments by reading their summaries and joining their results on `dataset_record_id`. The [CI action](/evals/ci) and the [compare cookbook](/evals/cookbooks/compare-versions) both do this for you.
* **An experiment does not pin an agent version.** Record the version you tested in the experiment `config`, or pin each input with a [context manifest](/evals/offline#pin-the-context-that-changed).

## Names you will see in code

| Term | Meaning |
| - | - |
| `task` | The function under test. It receives a case's `input` (and optionally hooks) and returns the output. |
| `case_id` | A case's stable ID within one experiment. Dataset-backed runs use the record ID; inline cases use your `id`, or `case-<n>`. |
| `subject_kind`, `subject_id` | The agent, harness, or skill the experiment evaluates. When set, the run appears on that agent's profile in the Atlan app. |
| `baseline_experiment_id` | The experiment this run should be compared with. |
| `config` | Free-form run settings: model, prompt version, git SHA. It is pinned at creation and filterable with `?config=key:value`. |
