Eval is the shortest path, but you may already have a harness: a benchmark runner, a simulator, or a job in another language. You can keep it and record its runs as experiments. You get the same datasets, summaries, CI gate, and comparisons.
A run has four steps. Eval performs the same steps internally.
A complete runner
my_runner.py
- A scorer ID and version on every score. A score span without them is dropped. Create a
codescorer once and reuse it. Because the scorer has no definition stored in Atlan, give it a new name when its logic changes. trace_idon the result must be the 32-character lowercase hex ID of that case’s trace. Read it from the root span, as above. Do not invent one.dataset_record_idis required on every result of a record-backed experiment, and it must be a record in the pinned snapshot. For inline cases, with no dataset, omit it.case_idis unique within the experiment. A second result with the samecase_idis refused, which makes an upload retry safe.- Flush before you finalize. Traces are exported in the background.
logger.flush()sends them, and Atlan indexes them a few seconds later.
CSV-backed datasets
A dataset created from a CSV file pins the file rather than records. Read the pinned version’s rows, run them, and name each result by itscase_id column. Do not set dataset_record_id.
case_id must exist in the pinned CSV. The CI action and comparisons pair CSV-backed runs on case_id.
From another language
The same four steps are plain HTTP. SendAuthorization: Bearer <token> and X-Atlan-Workspace-Id on every call, and include workspace_id in each create body.
scores values must be numbers or null. A result for a completed or failed experiment is refused with 409. To make scores visible in trace statistics as well, emit a score span per score. It is a child span with atlan.span.type = "score" and the attributes atlan.score.name, atlan.score.value, atlan.score.scorer_id, and atlan.score.scorer_version.