Skip to main content

Eval options

Python: Eval(name, *, data | dataset, task, scores, **options), or await async_eval(...) with the same arguments. TypeScript: await Eval(name, { data | dataset, task, scores, experimentName?, description? }, options). Task hooks (the second task parameter): expected, metadata, tags, span (the task span; call .update(...) on it), and trial_index / trialIndex, which is always 0.

Return value

scores is computed by Atlan from the stored results when the experiment finalizes. null scores are excluded from mean and count. metrics is written by the SDK, with snake_case keys from Python and camelCase from TypeScript (successfulCases, failedCases, durationMs). Cost and tokens are not in the summary; query them with trace stats.

Environment variables

AtlanClient reads no environment variables. Pass the origin, token, and workspace explicitly, as the examples do.

Statuses

Limits

Error codes

API routes

All under /eval/v1. Send Authorization: Bearer <token> and X-Atlan-Workspace-Id. Create bodies need workspace_id. PATCH is the update verb. Useful list filters: experiments take dataset_id, subject_kind and subject_id, experiment_status, baseline_experiment_id, config=key:value, and sort=-created_at. Results take case_id, dataset_record_id, and session_id. Records take q, label, category, source_kind, and source_ref. The SDK exposes the same tree: client.datasets, client.datasets.records, client.experiments, client.experiments.results, client.experiments.traces, and client.scorers. See the generated datasets, experiments, and scorers references.