Skip to main content
A dataset holds the cases every run is measured on. Inline cases are fine while you iterate. Move to a dataset when you need two runs, or a PR and its baseline, to use exactly the same cases.

Inline cases or a dataset

An inline case is a mapping with input (any JSON value) and optional expected, metadata, tags, and id. Set id on every inline case you may want to resume or compare.

Push a dataset from code

push_dataset in Python and pushDataset in TypeScript create the dataset once and sync its records on every call. A record whose content has not changed is left alone, so pushing the same suite on every CI run creates no empty versions.
Then run against it by ID or exact name. A name must match exactly one dataset in the workspace. A missing or ambiguous name fails rather than picking a close match.
The dataset name is unique in the workspace, and pushing to a name that already exists adds to that dataset. Check suite.created, and each record’s action in suite.records, if the name might already be taken. Each entry of suite.records also maps your case name (key) to its record ID (id); results carry only the record ID as case_id, so keep the map to print readable failures: names = {r.id: r.key for r in suite.records}.
A push never deletes records. A case you remove from your code stays in the dataset and keeps running. Archive it in the Atlan app, or with client.datasets.records.archive(dataset_id, record_id).

How a record reaches your task and scorer

If the readable question is not under question, set input_key when the dataset is created: push_dataset(client, name, records, input_key="prompt"). It cannot be changed by a later push. Atlan exposes input[input_key] as display_input on each record, so tables in the app and coding agents know which field is the question. Pass the whole input object instead of one field by keeping it under the key. For example, {"question": {"text": "...", "locale": "fr"}} hands your task the inner object.

Versions and snapshots

Editing a record appends a new immutable version; the old one stays readable at GET /eval/v1/datasets/{dataset_id}/records/{record_id}/versions/{n}. When an experiment starts, Atlan stamps a dataset_snapshot with every live record’s ID, version, and content hash, and Eval runs from that snapshot. A correction therefore changes only later runs. A snapshot holds at most 10,000 records. Split a larger suite into several datasets.

Record where a case came from

source_kind and source_ref record provenance. Set them when the record is first pushed; a later push does not change them. Use label for curation state such as needs_review, and categories for the scenario.

Filter records

GET /eval/v1/datasets/{id}/records accepts q (a substring over input, expected, label, and categories), label, category, source_kind, and source_ref. The SDK exposes the same filters on client.datasets.records.list.

CSV datasets

A dataset can also be backed by an uploaded CSV file, for example one created in the Atlan app. The CSV needs a unique, non-empty case_id column and a column named after the dataset’s input_key. Files are limited to 32 MiB, 100,000 rows, and 256 columns. Each upload is an immutable dataset version, and its content is readable at GET /eval/v1/datasets/{id}/versions/{n}/content. To create one over the API, reserve an upload, send the bytes, then create the dataset with the upload key:
Eval runs record-backed datasets, the kind push_dataset creates. To run a CSV-backed dataset, read its content and record results with your own runner. Each result names its row by case_id instead of dataset_record_id. See bring your own runner.