> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Bring your own runner

> Record an existing harness's runs as Atlan experiments: start the experiment, trace each case, upload results, and finalize, including CSV-backed datasets.

`Eval` is the shortest path, but you may already have a harness: a benchmark runner, a simulator, or a job in another language. You can keep it and record its runs as experiments. You get the same datasets, summaries, CI gate, and comparisons.

A run has four steps. `Eval` performs the same steps internally.

| Step | Call | Rule |
| - | - | - |
| 1. Start | `start_experiment` / `startExperiment`, or `POST /eval/v1/experiments` | Names are unique per workspace. The experiment starts `running` and pins the dataset snapshot. |
| 2. Trace each case | Run it inside `run.trace()` (Python) or `propagateAttributes(run.traceOptions, fn)` (TypeScript) | Every span is stamped with the experiment ID, so it appears in the experiment's trace views. Use one root span per case. |
| 3. Upload results | `experiments.results.create_bulk`, or `POST /eval/v1/experiments/{id}/results/bulk` | At most 100 per call. Each item is accepted or refused on its own: check every status. Results are create-only. |
| 4. Finalize | `experiments.update(id, {"experiment_status": "completed"})` | Atlan computes `summary.scores` from the results. The experiment is then frozen. Use `failed` for a run you abandon. |

## A complete runner

```python my_runner.py theme={null}
import os
import time
from atlanai import AtlanClient, start_experiment
from atlanai.tracing import init_logger

WORKSPACE = os.environ["ATLAN_WORKSPACE_ID"]
client = AtlanClient(
    os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
    bearer_token=os.environ["ATLAN_API_KEY"],
    workspace=WORKSPACE,
)
logger = init_logger(project_name="my-harness")
tracer = logger.client


def my_harness(question: str) -> str:   # your existing runner
    return "processed" if "refund" in question else "Friday"


# 1. A code scorer, created once. Its id and version go on every score.
found = client.scorers.list(limit=500).items
scorer = next((s for s in found if s.name == "my-harness-exact-match"), None) or client.scorers.create(
    {"workspace_id": WORKSPACE, "name": "my-harness-exact-match", "scorer_kind": "code"})

# 2. Start the experiment. Atlan pins the dataset snapshot.
run = start_experiment(client, "support-agent-cases",
                       {"name": f"my-harness-{int(time.time())}", "config": {"runner": "my-harness", "model": "example-model"}})

# 3. Run each case in its own trace, inside the experiment's scope.
items = []
for record in run.experiment.dataset_snapshot.records:
    question, expected = record.input["question"], record.expected["value"]
    with run.trace(), tracer.start_as_current_span("case", as_type="task") as root:
        trace_id = root.trace_id
        root.update(input=question)
        output = my_harness(question)
        root.update(output=output)
        score = float(output == expected)
        root.score_trace("exact_match", value=score,
                         scorer_id=scorer.id, scorer_version=scorer.version_ordinal)
    items.append({
        "workspace_id": WORKSPACE, "name": f"{run.id}-{record.id}",
        "case_id": record.id, "dataset_record_id": record.id,
        "input": record.input, "expected": record.expected, "output": {"value": output},
        "scores": {"exact_match": score}, "trace_id": trace_id,
    })

# 4. Export the traces, upload results (100 per call), then finalize.
logger.flush()
for i in range(0, len(items), 100):
    report = client.experiments.results.create_bulk(run.id, {"items": items[i:i + 100]})
    failed = [r for r in report.items if r.status_code >= 300]
    if failed:
        raise SystemExit(f"{len(failed)} results were refused: {failed[0].error}")

done = client.experiments.update(run.id, {"experiment_status": "completed",
                                          "summary": {"metrics": {"cases": len(items)}}})
print(done.id, done.summary["scores"])
```

What each rule protects:

* **A scorer ID and version on every score.** A score span without them is dropped. Create a `code` scorer once and reuse it. Because the scorer has no definition stored in Atlan, give it a new name when its logic changes.
* **`trace_id` on the result** must be the 32-character lowercase hex ID of that case's trace. Read it from the root span, as above. Do not invent one.
* **`dataset_record_id`** is required on every result of a record-backed experiment, and it must be a record in the pinned snapshot. For inline cases, with no dataset, omit it.
* **`case_id`** is unique within the experiment. A second result with the same `case_id` is refused, which makes an upload retry safe.
* **Flush before you finalize.** Traces are exported in the background. `logger.flush()` sends them, and Atlan indexes them a few seconds later.

After a few seconds, confirm that the evidence is complete:

```python theme={null}
from atlanai import verify_experiment

verify_experiment(client, run.id).raise_for_status()
```

## CSV-backed datasets

A dataset created from a CSV file pins the file rather than records. Read the pinned version's rows, run them, and name each result by its `case_id` column. Do not set `dataset_record_id`.

```python theme={null}
import csv
import io
import urllib.request

snapshot = run.experiment.dataset_snapshot          # dataset_id, version_ordinal, file pin
url = (f"{os.environ.get('ATLAN_BASE_URL', 'https://api.atlan.com')}"
       f"/eval/v1/datasets/{snapshot.dataset_id}/versions/{snapshot.version_ordinal}/content")
request = urllib.request.Request(url, headers={
    "Authorization": f"Bearer {os.environ['ATLAN_API_KEY']}",
    "X-Atlan-Workspace-Id": WORKSPACE,
})
rows = list(csv.DictReader(io.StringIO(urllib.request.urlopen(request).read().decode())))

for row in rows:
    ...  # run and trace the case as above, then:
    items.append({"workspace_id": WORKSPACE, "name": f"{run.id}-{row['case_id']}",
                  "case_id": row["case_id"], "input": {"question": row["question"]},
                  "expected": {"value": row["expected"]}, "output": {"value": output},
                  "scores": {"exact_match": score}, "trace_id": trace_id})
```

Every `case_id` must exist in the pinned CSV. The CI action and comparisons pair CSV-backed runs on `case_id`.

## From another language

The same four steps are plain HTTP. Send `Authorization: Bearer <token>` and `X-Atlan-Workspace-Id` on every call, and include `workspace_id` in each create body.

```text theme={null}
POST  /eval/v1/experiments                      {workspace_id, name, dataset_id, config}      → 201, experiment_status "running"
      ... run each case; export its spans over OTLP to /otel/v1/traces with the attribute
          atlan.eval.experiment_id = <experiment id> on every span
POST  /eval/v1/experiments/{id}/results/bulk    {items: [{workspace_id, name, case_id, dataset_record_id,
                                                 input, expected, output, scores, trace_id, duration_ms, error}]}
                                               → 207, one {index, status_code, record | error} per item
PATCH /eval/v1/experiments/{id}                 {experiment_status: "completed", summary: {metrics: {...}}}
                                               → 200, with summary.scores computed
```

`scores` values must be numbers or `null`. A result for a `completed` or `failed` experiment is refused with `409`. To make scores visible in trace statistics as well, emit a score span per score. It is a child span with `atlan.span.type = "score"` and the attributes `atlan.score.name`, `atlan.score.value`, `atlan.score.scorer_id`, and `atlan.score.scorer_version`.
