Inline cases or a dataset
An inline case is a mapping with
input (any JSON value) and optional expected, metadata, tags, and id. Set id on every inline case you may want to resume or compare.
Push a dataset from code
push_dataset in Python and pushDataset in TypeScript create the dataset once and sync its records on every call. A record whose content has not changed is left alone, so pushing the same suite on every CI run creates no empty versions.
suite.created, and each record’s action in suite.records, if the name might already be taken. Each entry of suite.records also maps your case name (key) to its record ID (id); results carry only the record ID as case_id, so keep the map to print readable failures: names = {r.id: r.key for r in suite.records}.
How a record reaches your task and scorer
If the readable question is not under
question, set input_key when the dataset is created: push_dataset(client, name, records, input_key="prompt"). It cannot be changed by a later push. Atlan exposes input[input_key] as display_input on each record, so tables in the app and coding agents know which field is the question.
Pass the whole input object instead of one field by keeping it under the key. For example, {"question": {"text": "...", "locale": "fr"}} hands your task the inner object.
Versions and snapshots
Editing a record appends a new immutable version; the old one stays readable atGET /eval/v1/datasets/{dataset_id}/records/{record_id}/versions/{n}. When an experiment starts, Atlan stamps a dataset_snapshot with every live record’s ID, version, and content hash, and Eval runs from that snapshot. A correction therefore changes only later runs.
A snapshot holds at most 10,000 records. Split a larger suite into several datasets.
Record where a case came from
source_kind and source_ref record provenance. Set them when the record is first pushed; a later push does not change them.
Use
label for curation state such as needs_review, and categories for the scenario.
Filter records
GET /eval/v1/datasets/{id}/records accepts q (a substring over input, expected, label, and categories), label, category, source_kind, and source_ref. The SDK exposes the same filters on client.datasets.records.list.
CSV datasets
A dataset can also be backed by an uploaded CSV file, for example one created in the Atlan app. The CSV needs a unique, non-emptycase_id column and a column named after the dataset’s input_key. Files are limited to 32 MiB, 100,000 rows, and 256 columns. Each upload is an immutable dataset version, and its content is readable at GET /eval/v1/datasets/{id}/versions/{n}/content.
To create one over the API, reserve an upload, send the bytes, then create the dataset with the upload key:
Eval runs record-backed datasets, the kind push_dataset creates. To run a CSV-backed dataset, read its content and record results with your own runner. Each result names its row by case_id instead of dataset_record_id. See bring your own runner.