> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Datasets

> Keep eval cases in a versioned dataset, map records to your task's input and your scorer's expected value, and track where each case came from.

A dataset holds the cases every run is measured on. Inline cases are fine while you iterate. Move to a dataset when you need two runs, or a PR and its baseline, to use exactly the same cases.

## Inline cases or a dataset

| | Inline `data` | Registry dataset |
| - | - | - |
| Where it lives | In your eval code | In the workspace, visible under **Evals → Datasets** |
| Case identity | Your `id`, or `case-<n>` | The record ID, stable across runs |
| Comparison across runs | Only if you keep the `id`s stable | Built in: results join on `dataset_record_id` |
| Snapshot at run start | A digest of the cases, used for resume | A full `dataset_snapshot` with every record's version and hash |
| Use for | Local iteration, quick checks | Baselines, CI gates, model comparisons |

An inline case is a mapping with `input` (any JSON value) and optional `expected`, `metadata`, `tags`, and `id`. Set `id` on every inline case you may want to resume or compare.

## Push a dataset from code

`push_dataset` in Python and `pushDataset` in TypeScript create the dataset once and sync its records on every call. A record whose content has not changed is left alone, so pushing the same suite on every CI run creates no empty versions.

<CodeGroup>
  ```python Python theme={null}
  import os
  from atlanai import AtlanClient, push_dataset

  client = AtlanClient(
      os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
      bearer_token=os.environ["ATLAN_API_KEY"],
      workspace=os.environ["ATLAN_WORKSPACE_ID"],
  )

  suite = push_dataset(client, "support-agent-cases", [
      {
          "name": "refund-status",                       # your stable key for the case
          "input": {"question": "Where is my refund for order 1042?"},
          "expected": {"value": "processed"},
          "categories": ["refunds"],
          "extra": {"priority": "high"},
      },
      {
          "name": "shipping-eta",
          "input": {"question": "When will order 2210 arrive?"},
          "expected": {"value": "Friday"},
          "categories": ["shipping"],
      },
  ])
  print(suite.id, [r.action for r in suite.records])
  # First push: ['created', 'created']. Later pushes: 'unchanged' or 'updated' per record.
  ```

  ```typescript TypeScript theme={null}
  import { AtlanClient, pushDataset } from "@atlanai/sdk";

  const client = new AtlanClient({
    gatewayOrigin: process.env.ATLAN_BASE_URL ?? "https://api.atlan.com",
    bearerToken: process.env.ATLAN_API_KEY!,
    workspace: process.env.ATLAN_WORKSPACE_ID!,
  });

  const suite = await pushDataset(client, "support-agent-cases", [
    {
      name: "refund-status",
      input: { question: "Where is my refund for order 1042?" },
      expected: { value: "processed" },
      categories: ["refunds"],
      extra: { priority: "high" },
    },
    {
      name: "shipping-eta",
      input: { question: "When will order 2210 arrive?" },
      expected: { value: "Friday" },
      categories: ["shipping"],
    },
  ]);
  console.log(suite.id);
  ```
</CodeGroup>

Then run against it by ID or exact name. A name must match exactly one dataset in the workspace. A missing or ambiguous name fails rather than picking a close match.

```python theme={null}
from atlanai import Eval

run = Eval("support-agent", dataset=suite.id, task=answer, scores=[accuracy])
```

The dataset name is unique in the workspace, and pushing to a name that already exists **adds to that dataset**. Check `suite.created`, and each record's `action` in `suite.records`, if the name might already be taken. Each entry of `suite.records` also maps your case name (`key`) to its record ID (`id`); results carry only the record ID as `case_id`, so keep the map to print readable failures: `names = {r.id: r.key for r in suite.records}`.

<Warning>
  A push never deletes records. A case you remove from your code stays in the dataset and keeps running. Archive it in the Atlan app, or with `client.datasets.records.archive(dataset_id, record_id)`.
</Warning>

## How a record reaches your task and scorer

| Record field | Becomes | Notes |
| - | - | - |
| `input[input_key]` | The value passed to `task` | `input` must be a JSON object. `input_key` defaults to `question`. A record without that key fails the run. |
| `expected` | `expected` in your scorer | A single-key `{"value": x}` is unwrapped to `x`. Any other shape is passed as stored. |
| `extra` | `metadata` in your scorer and in task hooks | Use it for case-specific settings, such as a fixture ID or a tolerance. |
| `categories` | The case's `tags` | Use them to slice results: which categories regressed? |
| Record ID | `case_id` and `dataset_record_id` on the result | The join key across experiments. |

If the readable question is not under `question`, set `input_key` when the dataset is created: `push_dataset(client, name, records, input_key="prompt")`. It cannot be changed by a later push. Atlan exposes `input[input_key]` as `display_input` on each record, so tables in the app and coding agents know which field is the question.

Pass the whole input object instead of one field by keeping it under the key. For example, `{"question": {"text": "...", "locale": "fr"}}` hands your task the inner object.

## Versions and snapshots

Editing a record appends a new immutable version; the old one stays readable at `GET /eval/v1/datasets/{dataset_id}/records/{record_id}/versions/{n}`. When an experiment starts, Atlan stamps a `dataset_snapshot` with every live record's ID, version, and content hash, and `Eval` runs from that snapshot. A correction therefore changes only later runs.

A snapshot holds at most 10,000 records. Split a larger suite into several datasets.

## Record where a case came from

`source_kind` and `source_ref` record provenance. Set them when the record is first pushed; a later push does not change them.

| `source_kind` | `source_ref` | Use |
| - | - | - |
| `manual` (default) | none | Written by hand |
| `session` | A session ID in the same workspace; Atlan checks that it exists | Curated from a production session. See [turn a failure into a test](/evals/cookbooks/production-to-dataset). |
| `trace` | A trace ID | Curated from one trace |
| `external` | Any reference up to 255 characters | Imported from another tool or a ticket |

Use `label` for curation state such as `needs_review`, and `categories` for the scenario.

## Filter records

`GET /eval/v1/datasets/{id}/records` accepts `q` (a substring over input, expected, label, and categories), `label`, `category`, `source_kind`, and `source_ref`. The SDK exposes the same filters on `client.datasets.records.list`.

## CSV datasets

A dataset can also be backed by an uploaded CSV file, for example one created in the Atlan app. The CSV needs a unique, non-empty `case_id` column and a column named after the dataset's `input_key`. Files are limited to 32 MiB, 100,000 rows, and 256 columns. Each upload is an immutable dataset version, and its content is readable at `GET /eval/v1/datasets/{id}/versions/{n}/content`.

To create one over the API, reserve an upload, send the bytes, then create the dataset with the upload key:

```bash theme={null}
# 1. Reserve an upload; the response carries upload_key.
curl -sS --fail-with-body -X POST "${ATLAN_BASE_URL}/file/v1/files/uploads" \
  -H "Authorization: Bearer ${ATLAN_API_KEY}" -H "Content-Type: application/json" \
  -d "{\"size_bytes\": $(wc -c < cases.csv)}"

# 2. Send the file.
curl -sS --fail-with-body -X PUT "${ATLAN_BASE_URL}/file/v1/files/uploads/${UPLOAD_KEY}" \
  -H "Authorization: Bearer ${ATLAN_API_KEY}" -H "Content-Type: application/octet-stream" \
  --data-binary @cases.csv

# 3. Create the dataset. Use POST /eval/v1/datasets/{id}/versions to add a version.
curl -sS --fail-with-body -X POST "${ATLAN_BASE_URL}/eval/v1/datasets" \
  -H "Authorization: Bearer ${ATLAN_API_KEY}" -H "Content-Type: application/json" \
  -H "X-Atlan-Workspace-Id: ${ATLAN_WORKSPACE_ID}" \
  -d "{\"workspace_id\": \"${ATLAN_WORKSPACE_ID}\", \"name\": \"csv-cases\", \"upload_key\": \"${UPLOAD_KEY}\", \"filename\": \"cases.csv\"}"
```

<Note>
  `Eval` runs record-backed datasets, the kind `push_dataset` creates. To run a CSV-backed dataset, read its content and record results with your own runner. Each result names its row by `case_id` instead of `dataset_record_id`. See [bring your own runner](/evals/custom-runner#csv-backed-datasets).
</Note>
