> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Evals reference

> Every Eval option, the result and summary shapes, environment variables, limits, statuses, error codes, and the eval API routes.

## Eval options

Python: `Eval(name, *, data | dataset, task, scores, **options)`, or `await async_eval(...)` with the same arguments. TypeScript: `await Eval(name, { data | dataset, task, scores, experimentName?, description? }, options)`.

| Python | TypeScript | Default | Meaning |
| - | - | - | - |
| `name` | `name` | required | Stable eval name. Auto-registered scorers are named `<name>-<scorer>`. The CI action matches baselines on it. |
| `data` | `data` | one of `data`/`dataset` | Inline cases: a list, iterable, async iterable, or a callable returning one. Each case has `input` and optional `expected`, `metadata`, `tags`, and `id`. |
| `dataset` | `dataset` | one of `data`/`dataset` | A dataset ID or its exact name. A record-backed dataset only. |
| `task` | `task` | required | `task(input)` or `task(input, hooks)`, sync or async |
| `scores` | `scores` | required, 1 or more | Scorers; see [the contract](/evals/scorers#the-contract) |
| `experiment_name` | `experimentName` | `name` | Display name of the experiment |
| `description` | `description` | none | Experiment description |
| `config` | `config` | `{}` | Free-form run settings, stored on the experiment and filterable |
| `baseline_experiment_id` | `baselineExperimentId` | none | The experiment to compare with |
| `subject_kind`, `subject_id` | `subjectKind`, `subjectId` | none | `agent`, `harness`, or `skill`, and its ID. Set both or neither. |
| `context_manifest` | `contextManifest` | none | A [context manifest](/evals/offline#pin-the-context-that-changed) |
| `max_concurrency` | not available | `1` | Cases in flight (Python). TypeScript runs cases one at a time. |
| `resume_experiment_id` | `resumeExperimentId` | none | Continue a `running` experiment |
| `on_start` | `onStart` | none | Called with the experiment ID when the run starts |
| `client` | `client` | from the environment | A configured `AtlanClient` |
| `trace_logger` | `logger` | a new logger | Your tracing logger. It must export every span with content. |
| `gateway_origin`, `api_key`, `workspace_id`, `project_id` | `gatewayOrigin`, `apiKey`, `workspaceId`, `projectId` | from the environment | Connection overrides |
| `return_results` | `returnResults` | `true` | Include per-case results in the return value |
| `verify_traces` | `verifyTraces` | `true` | Wait until every case trace is queryable before uploading its result |
| `trace_verification_timeout_s` | `traceVerificationTimeoutMs` | `30` s | How long to wait for each trace |

**Task hooks** (the second task parameter): `expected`, `metadata`, `tags`, `span` (the task span; call `.update(...)` on it), and `trial_index` / `trialIndex`, which is always `0`.

## Return value

| Python | TypeScript | Contents |
| - | - | - |
| `experiment_id` | `experimentId` | The experiment ID |
| `experiment` | `experiment` | The finalized experiment, including `dataset_snapshot` |
| `summary` | `summary` | See below |
| `results` | `results` | One per case: `case_id`/`caseId`, `input`, `expected`, `metadata`, `tags`, `dataset_record_id`, `output`, `error`, `scores`, `trace_id`/`traceId`, `duration_ms`/`durationMs` |
| `dataset` | `dataset` | The resolved dataset, when one was used |

```json theme={null}
{
  "scores":  {"accuracy": {"mean": 0.5, "count": 2}},
  "scorers": {"accuracy": {"scorer_id": "scorer_01example", "scorer_version": 1}},
  "metrics": {"cases": 2, "successful_cases": 2, "failed_cases": 0, "duration_ms": 3101}
}
```

`scores` is computed by Atlan from the stored results when the experiment finalizes. `null` scores are excluded from `mean` and `count`. `metrics` is written by the SDK, with snake\_case keys from Python and camelCase from TypeScript (`successfulCases`, `failedCases`, `durationMs`). Cost and tokens are not in the summary; query them with [trace stats](/evals/agents#cost-tokens-and-latency).

## Environment variables

| Variable | Used by | Meaning |
| - | - | - |
| `ATLAN_API_KEY` | `Eval`, tracing | Bearer token for API calls and trace export |
| `ATLAN_WORKSPACE_ID` | `Eval`, tracing | Workspace for every object the run creates |
| `ATLAN_BASE_URL` | `Eval`, tracing | Gateway origin. Default `https://api.atlan.com`. |
| `ATLAN_DEBUG` | tracing | `true` logs dropped-score and export warnings |
| `ATLAN_TRACING`, `ATLAN_TRACE_CONTENT` | tracing | Either set to `false` makes `Eval` refuse to run, because results would lack evidence |
| `ATLAN_EVAL_GIT_BRANCH`, `ATLAN_EVAL_GIT_SHA`, `ATLAN_EVAL_GIT_PR`, `ATLAN_EVAL_MODEL`, `ATLAN_EVAL_RESULTS_FILE` | Python `Eval`, and the [TypeScript shim](/evals/ci#typescript-evals) | Set by the [CI action](/evals/ci) for any command. Set them yourself only to simulate the action locally. |

`AtlanClient` reads no environment variables. Pass the origin, token, and workspace explicitly, as the examples do.

## Statuses

| Object | Values |
| - | - |
| Experiment `experiment_status` | `running` → `completed` or `failed`. A terminal experiment is frozen. |
| Scoring run `run_status` | `queued` → `running` → `succeeded` or `failed` |
| Record `source_kind` | `manual`, `session`, `trace`, `external` |
| Scorer `scorer_kind` | `code`, `llm_judge`, `human` |

## Limits

| Limit | Value |
| - | - |
| Records in a dataset snapshot | 10,000 |
| Results folded into a summary | 10,000 per experiment |
| Items per bulk create (records or results) | 100 |
| List page size | default 50, maximum 500 |
| CSV dataset | 32 MiB, 100,000 rows, 256 columns |
| `case_id` | Unique per experiment, up to 255 characters |
| `trace_id` on a result | 32 lowercase hex characters, not all zeros |
| Score comment from `ScoreValue` metadata | 2,000 characters |
| Hosted scoring | 100 sessions per run, 5 in `preview`/`test`; default concurrency 4 |

## Error codes

| Status | Typical cause |
| - | - |
| `400` | A validation failure: a bad `trace_id`, a missing `dataset_record_id` on a record-backed run, a `case_id` not in the pinned CSV, a half-set subject pair, or an invalid scorer definition or selector |
| `403` | Scoring or archiving without the required role |
| `404` | Not found, not visible to you, or a child requested under the wrong parent |
| `409` | Writing to a `completed` or `failed` experiment, a duplicate `case_id` or name, or an unknown workspace |
| `412` | A stale `If-Match` on a scorer edit |
| `422` | An unknown enum value, or moving a dataset or experiment to another workspace |
| `503` | Traces are temporarily unreadable, or no judge backend is configured for hosted scoring |
| `207` | A bulk create: read each item's `status_code` |

## API routes

All under `/eval/v1`. Send `Authorization: Bearer <token>` and `X-Atlan-Workspace-Id`. Create bodies need `workspace_id`. `PATCH` is the update verb.

| Resource | Routes |
| - | - |
| Datasets | `POST /datasets`, `GET /datasets`, `GET·PATCH /datasets/{id}`, `POST /datasets/{id}/archive`, `POST /datasets/search` |
| Dataset versions (CSV) | `POST·GET /datasets/{id}/versions`, `GET /datasets/{id}/versions/{n}`, `GET /datasets/{id}/versions/{n}/content` |
| Records | `POST /datasets/{id}/records`, `POST /datasets/{id}/records/bulk`, `GET /datasets/{id}/records`, `GET·PATCH /datasets/{id}/records/{rid}`, `POST …/records/{rid}/archive`, `GET …/records/{rid}/versions`, `GET …/records/{rid}/versions/{n}` |
| Experiments | `POST /experiments`, `GET /experiments`, `GET·PATCH /experiments/{id}`, `POST /experiments/{id}/archive`, `POST /experiments/search` |
| Results | `POST /experiments/{id}/results`, `POST /experiments/{id}/results/bulk`, `GET /experiments/{id}/results`, `GET /experiments/{id}/results/{rid}` |
| Experiment traces | `GET /experiments/{id}/traces`, `GET /experiments/{id}/traces/{tid}`, `GET /experiments/{id}/traces/{tid}/spans`, `POST /experiments/{id}/traces/stats/query` |
| Scorers | `POST /scorers`, `GET /scorers`, `GET·PATCH /scorers/{id}`, `GET /scorers/{id}/versions`, `GET /scorers/{id}/versions/{n}`, `POST /scorers/{id}/archive`, `GET /scorer-starters`, `POST /scorers/{id}/score` |

Useful list filters: experiments take `dataset_id`, `subject_kind` and `subject_id`, `experiment_status`, `baseline_experiment_id`, `config=key:value`, and `sort=-created_at`. Results take `case_id`, `dataset_record_id`, and `session_id`. Records take `q`, `label`, `category`, `source_kind`, and `source_ref`.

The SDK exposes the same tree: `client.datasets`, `client.datasets.records`, `client.experiments`, `client.experiments.results`, `client.experiments.traces`, and `client.scorers`. See the generated [datasets](/tools/sdk/references/operations/datasets), [experiments](/tools/sdk/references/operations/experiments), and [scorers](/tools/sdk/references/operations/scorers) references.
