> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Troubleshooting evals

> Symptoms, causes, and fixes for eval runs, datasets, scorers, CI gates, and hosted scoring.

Find the symptom, then apply the fix. Most issues show up in one of three places: the exception `Eval` raises, a result's `error` field, or the CI verdict.

## Running an eval

| Symptom | Cause | Fix |
| - | - | - |
| `... reuses a Registry name with a different identity` on the second run of an eval | Your SDK is older than the Atlan gateway it talks to | Upgrade: `pip install -U 'atlanai[tracing]'` or `npm install @atlanai/sdk@latest` |
| `Eval() cannot run inside an event loop` | Python `Eval` was called from Jupyter, an async framework, or an async test | `await async_eval(...)` with the same arguments |
| `summary.scores` is `{}` | Every case's task raised, so no scorer ran | Read `run.results[i].error`. A `NameError` or `AttributeError` there is a bug in your task, not in the agent. |
| A score is missing from some cases | The scorer returned `None` (not applicable) or raised on those cases | A raise appears in that case's `error` as `"<scorer>: <message>"` |
| `... has no records, so there is nothing to evaluate` | The dataset is empty, or it is CSV-backed | Push records with `push_dataset`, or [run a CSV dataset with your own runner](/evals/custom-runner#csv-backed-datasets). The experiment the call created is left `running`. |
| `ValueError` naming `input_key` | A record's `input` lacks the dataset's `input_key` (default `question`) | Store the question under that key, or create the dataset with `input_key="..."` |
| The run raised and the experiment stays `running` | A trace did not arrive in time, a result upload was refused, or the process died | [Resume it](/evals/offline#resume-a-stopped-run) with the ID in the error message. For slow trace export, raise `trace_verification_timeout_s`. If you will not resume it, [close it](/evals/offline#close-a-run-you-will-not-resume). |
| `push_dataset` fails with `500` `record_push_failed` | A transient write failure | Retry the call. `push_dataset` is idempotent: records already stored come back `unchanged`. |
| `Eval` refuses to start, citing sampling or content | Your logger samples spans or drops content, or `ATLAN_TRACING`/`ATLAN_TRACE_CONTENT` is `false` | Eval traces must be complete. Pass a logger with sample rate `1`, or let `Eval` create one. |
| Two scorers "own" the same score name | Two functions return the same score name | Every score name belongs to one scorer. Rename one. |
| The baseline and candidate have different scorer IDs | They ran under different `Eval` names | Use one eval name; put the variant in `config` |
| `409` when you create an experiment | Experiment names are unique per workspace | `Eval` makes names unique for you. With `start_experiment`, add a timestamp or run ID. |

## Results and uploads (your own runner)

| Symptom | Cause | Fix |
| - | - | - |
| `400` on a result mentioning `trace_id` | It must be 32 lowercase hex characters, not all zeros | Read it from the root span: `root.trace_id` / `root.traceId` |
| `400` on a result mentioning `dataset_record_id` | Record-backed runs require it, and it must be in the pinned snapshot. Inline runs and CSV runs must omit it. | Use `experiment.dataset_snapshot.records[i].id`. For CSV, send `case_id` only. |
| `409` on a result | The experiment is already `completed` or `failed`, or the `case_id` was already uploaded | A finished experiment is frozen; start a new one. A duplicate `case_id` on retry is safe to ignore. |
| `207` with some items refused | Bulk creates validate each item | Read each item's `status_code` and `error`; do not assume the batch succeeded |
| `result.input` or `output` is `{"value": ...}` | Results store objects. A string or number is wrapped as `{"value": x}` | Unwrap it when you read results back |
| `verify_experiment` reports orphaned spans | The case span was opened with an invented `trace_context` | Let the SDK choose the trace ID and read it from the span |

## Scores from application code

| Symptom | Cause | Fix |
| - | - | - |
| A score never appears | It had no valid `scorer_id` and `scorer_version`, an empty name, or a value that does not match its `data_type` | Scores are dropped with a warning, never raised. Set `ATLAN_DEBUG=true` to see why. |
| A categorical score has no aggregate | Only numeric values are aggregated | Pass a numeric `value` alongside `string_value` if you need an average |

## CI

| Verdict or symptom | Cause | Fix |
| - | - | - |
| `invalid`: no experiments reported | The command did not call `Eval`, or a TypeScript eval did not write the results file | Use Python `Eval`, or add the [TypeScript shim](/evals/ci#typescript-evals) |
| `invalid`, and the command's log shows a missing `ATLAN_API_KEY` or workspace | The action does not pass its `api-key` input to the command | Set `ATLAN_API_KEY` and `ATLAN_WORKSPACE_ID` in the action step's `env` ([workflow](/evals/ci#2-add-the-workflow)) |
| A required eval check stays pending | A `paths:` filter skipped the workflow | Remove the filter; see the [note in the CI guide](/evals/ci#2-add-the-workflow) |
| `invalid`: no git provenance, or not from this commit or PR | `config.git` is missing or was set by hand | Let the SDK stamp it. Do not set `config.git` yourself, except in the TypeScript shim, which copies the action's values. |
| `invalid`: error rate | More than `max-error-rate` (default 20%) of cases errored | Fix the failing cases; the comment gives the count, and the experiment in Atlan gives the errors |
| `no-baseline` on every PR | No completed run on the default branch with the same eval name and dataset | Run the workflow on `main` once; keep the eval name identical across branches |
| A regression you cannot reproduce | Run-to-run variance is larger than the threshold | Raise `min-regressed-cases`, widen `thresholds`, add cases, or set temperature to `0` |
| The job has no API key on fork PRs | GitHub does not pass secrets to forks | Skip forks with the `if:` shown in the [workflow](/evals/ci#2-add-the-workflow) |

## Hosted scoring

| Symptom | Cause | Fix |
| - | - | - |
| `503` from `/score` | No judge backend is configured on the deployment | Ask your Atlan administrator; `preview` also requires a backend |
| `403` from `/score` | You need the builder or admin role in the **scorer's** workspace | Request the role, or create the scorer in a workspace where you have it |
| `400` naming a selector | Two selector families in one call, or none matched | Send one family: IDs, a time window, agents, or an experiment |
| A session line says there is nothing to judge | The session has no readable trace | Check that the trace is ingested and belongs to the scorer's workspace |
| A line says the model gateway returned `404` | The deployment's judge model is not registered in your organization's model catalog | Ask your administrator to register it |
| A line fails on context length | The transcript is larger than the judge model's context. Transcripts are not truncated. | Use `grain: "trace"` rather than `session`, or narrow `state` to the fields that matter |
| The `experiment_id` selector matches nothing | Results written by `Eval` carry no session | Select the sessions by ID or time window instead |
| New sessions are not scored by **Run live** | Only sessions that reach `completed` or `failed` are scored automatically | Make sure your agent records the session's end, or score on demand |
| `confidence` looks low for a confident answer | `confidence` is chance-corrected, not a probability | Read `probabilities[choice]`; see [reading an answer](/evals/judges#reading-an-answer) |
