Eval raises, a result’s error field, or the CI verdict.
Running an eval
| Symptom | Cause | Fix |
|---|---|---|
... reuses a Registry name with a different identity on the second run of an eval | Your SDK is older than the Atlan gateway it talks to | Upgrade: pip install -U 'atlanai[tracing]' or npm install @atlanai/sdk@latest |
Eval() cannot run inside an event loop | Python Eval was called from Jupyter, an async framework, or an async test | await async_eval(...) with the same arguments |
summary.scores is {} | Every case’s task raised, so no scorer ran | Read run.results[i].error. A NameError or AttributeError there is a bug in your task, not in the agent. |
| A score is missing from some cases | The scorer returned None (not applicable) or raised on those cases | A raise appears in that case’s error as "<scorer>: <message>" |
... has no records, so there is nothing to evaluate | The dataset is empty, or it is CSV-backed | Push records with push_dataset, or run a CSV dataset with your own runner. The experiment the call created is left running. |
ValueError naming input_key | A record’s input lacks the dataset’s input_key (default question) | Store the question under that key, or create the dataset with input_key="..." |
The run raised and the experiment stays running | A trace did not arrive in time, a result upload was refused, or the process died | Resume it with the ID in the error message. For slow trace export, raise trace_verification_timeout_s. If you will not resume it, close it. |
push_dataset fails with 500 record_push_failed | A transient write failure | Retry the call. push_dataset is idempotent: records already stored come back unchanged. |
Eval refuses to start, citing sampling or content | Your logger samples spans or drops content, or ATLAN_TRACING/ATLAN_TRACE_CONTENT is false | Eval traces must be complete. Pass a logger with sample rate 1, or let Eval create one. |
| Two scorers “own” the same score name | Two functions return the same score name | Every score name belongs to one scorer. Rename one. |
| The baseline and candidate have different scorer IDs | They ran under different Eval names | Use one eval name; put the variant in config |
409 when you create an experiment | Experiment names are unique per workspace | Eval makes names unique for you. With start_experiment, add a timestamp or run ID. |
Results and uploads (your own runner)
| Symptom | Cause | Fix |
|---|---|---|
400 on a result mentioning trace_id | It must be 32 lowercase hex characters, not all zeros | Read it from the root span: root.trace_id / root.traceId |
400 on a result mentioning dataset_record_id | Record-backed runs require it, and it must be in the pinned snapshot. Inline runs and CSV runs must omit it. | Use experiment.dataset_snapshot.records[i].id. For CSV, send case_id only. |
409 on a result | The experiment is already completed or failed, or the case_id was already uploaded | A finished experiment is frozen; start a new one. A duplicate case_id on retry is safe to ignore. |
207 with some items refused | Bulk creates validate each item | Read each item’s status_code and error; do not assume the batch succeeded |
result.input or output is {"value": ...} | Results store objects. A string or number is wrapped as {"value": x} | Unwrap it when you read results back |
verify_experiment reports orphaned spans | The case span was opened with an invented trace_context | Let the SDK choose the trace ID and read it from the span |
Scores from application code
| Symptom | Cause | Fix |
|---|---|---|
| A score never appears | It had no valid scorer_id and scorer_version, an empty name, or a value that does not match its data_type | Scores are dropped with a warning, never raised. Set ATLAN_DEBUG=true to see why. |
| A categorical score has no aggregate | Only numeric values are aggregated | Pass a numeric value alongside string_value if you need an average |
CI
| Verdict or symptom | Cause | Fix |
|---|---|---|
invalid: no experiments reported | The command did not call Eval, or a TypeScript eval did not write the results file | Use Python Eval, or add the TypeScript shim |
invalid, and the command’s log shows a missing ATLAN_API_KEY or workspace | The action does not pass its api-key input to the command | Set ATLAN_API_KEY and ATLAN_WORKSPACE_ID in the action step’s env (workflow) |
| A required eval check stays pending | A paths: filter skipped the workflow | Remove the filter; see the note in the CI guide |
invalid: no git provenance, or not from this commit or PR | config.git is missing or was set by hand | Let the SDK stamp it. Do not set config.git yourself, except in the TypeScript shim, which copies the action’s values. |
invalid: error rate | More than max-error-rate (default 20%) of cases errored | Fix the failing cases; the comment gives the count, and the experiment in Atlan gives the errors |
no-baseline on every PR | No completed run on the default branch with the same eval name and dataset | Run the workflow on main once; keep the eval name identical across branches |
| A regression you cannot reproduce | Run-to-run variance is larger than the threshold | Raise min-regressed-cases, widen thresholds, add cases, or set temperature to 0 |
| The job has no API key on fork PRs | GitHub does not pass secrets to forks | Skip forks with the if: shown in the workflow |
Hosted scoring
| Symptom | Cause | Fix |
|---|---|---|
503 from /score | No judge backend is configured on the deployment | Ask your Atlan administrator; preview also requires a backend |
403 from /score | You need the builder or admin role in the scorer’s workspace | Request the role, or create the scorer in a workspace where you have it |
400 naming a selector | Two selector families in one call, or none matched | Send one family: IDs, a time window, agents, or an experiment |
| A session line says there is nothing to judge | The session has no readable trace | Check that the trace is ingested and belongs to the scorer’s workspace |
A line says the model gateway returned 404 | The deployment’s judge model is not registered in your organization’s model catalog | Ask your administrator to register it |
| A line fails on context length | The transcript is larger than the judge model’s context. Transcripts are not truncated. | Use grain: "trace" rather than session, or narrow state to the fields that matter |
The experiment_id selector matches nothing | Results written by Eval carry no session | Select the sessions by ID or time window instead |
| New sessions are not scored by Run live | Only sessions that reach completed or failed are scored automatically | Make sure your agent records the session’s end, or score on demand |
confidence looks low for a confident answer | confidence is chance-corrected, not a probability | Read probabilities[choice]; see reading an answer |