> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# For coding agents

> A brief to hand a coding agent that is setting up evals, the invariants it must respect, a task-to-page map, and a checkable definition of done.

This page is written for a coding agent, such as Claude Code, Codex, or Cursor, that has been asked to set up evals for a project. It is equally a checklist for a person.

## Load the docs

* Every page is Markdown at `<page>.md`, for example `https://platform.atlan.com/evals/quickstart.md`. `https://platform.atlan.com/llms.txt` lists every page with its description.
* Connect the docs MCP server so the agent can search current docs instead of relying on training data: `claude mcp add --transport http atlan-docs https://platform.atlan.com/mcp`. See [For AI agents](/for-ai-agents) for other clients.
* Read in this order: [quickstart](/evals/quickstart), [concepts](/evals/concepts), [scorers](/evals/scorers), [agents](/evals/agents), [CI](/evals/ci).

## A brief to paste

```text theme={null}
Set up Atlan evals for this repository's agent.

1. Find the agent's entry point: the function that takes a user request and returns the
   answer. Write a thin adapter `task(input) -> {"answer": ..., "tool_calls": [{"name", "args"}]}`
   (`toolCalls` in TypeScript). Record tool calls where the tools execute, not from the
   model's own account. Replace tools with side effects by recording stand-ins.
2. Create evals/cases.py with 10–30 cases covering the main behaviours and known failures.
   Each case: a stable name, input {"question": ...}, and expected (the answer facts and
   the tools that should be called). Push them with atlanai.push_dataset.
3. Write deterministic scorers first (exact or required facts, tool choice, tool
   arguments). Each returns 0..1, or None when it does not apply. Add an LLM-judge scorer
   only for what code cannot decide.
4. Create evals/run.py (with an empty evals/__init__.py) that pushes the dataset and calls
   atlanai.Eval with a fixed eval name. Run it locally with `python -m evals.run`; every case
   must finish without an error. For TypeScript, create evals/run.ts with pushDataset and
   Eval from @atlanai/sdk, run it with `npx tsx evals/run.ts`, and make it pass
   `tsc --strict`.
5. Add .github/workflows/atlan-evals.yml using atlanai/agent-registry-action/eval,
   triggered on pull_request and on push to main, plus a requirements.txt with the agent's
   dependencies (or the TypeScript workflow in the CI page). Set ATLAN_API_KEY and
   ATLAN_WORKSPACE_ID in the action step's env: the action does not pass its api-key
   input to the command. A TypeScript eval must also write the results file.
6. Do not hardcode credentials. Read ATLAN_API_KEY and ATLAN_WORKSPACE_ID from the
   environment; the workflow reads them from secrets.ATLAN_API_KEY and
   vars.ATLAN_WORKSPACE_ID.

Docs: https://platform.atlan.com/evals/for-coding-agents.md. Stop and ask if the agent has side
effects (sends email, writes to production) that an eval run would trigger.
```

## Invariants

These are the facts an agent most often gets wrong. Each one is enforced by the API or the SDK.

1. **Atlan does not run the agent.** `Eval` executes in the local process or CI job. There is no hosted runner to configure.
2. **Install the latest SDK.** `pip install -U 'atlanai[tracing]'` for Python, `npm install @atlanai/sdk@latest` for TypeScript. The `tracing` extra is required in Python.
3. **`Eval` reads `ATLAN_API_KEY`, `ATLAN_WORKSPACE_ID`, and `ATLAN_BASE_URL`.** `AtlanClient` reads nothing from the environment; pass its arguments explicitly.
4. **Keep the eval name fixed.** Scorer identity and CI baselines key on it. Put variants in `config`.
5. **Scorers return `0..1`, higher is better, or `None` to abstain.** Never return a free `1.0` for a case a rule does not cover.
6. **Results are create-only and finished experiments are frozen.** To fix a run, start a new one. `summary.scores` is computed by Atlan, not written by you.
7. **The task sees only `input`, and scorers see only what the task returns.** Return the tool calls from the task if scorers need them. Record them at the tool, not from the agent's self-report.
8. **Dataset records need an object `input`** with the question under `input_key` (default `question`). `{"value": x}` as `expected` reaches the scorer as `x`.
9. **Hosted `llm_judge` scorers score recorded sessions, not `Eval` cases.** For a judge inside `Eval`, call a model from a code scorer.
10. **Python `Eval` cannot run inside an event loop.** Use `await async_eval(...)` in notebooks and async code.
11. **The CI action reads experiments that Python `Eval` reports.** A TypeScript eval needs the [results-file shim](/evals/ci#typescript-evals). The action's `api-key` never reaches the command; set `ATLAN_API_KEY` and `ATLAN_WORKSPACE_ID` in the step's `env`.
12. **An agent version is not pinned automatically.** Record the commit and model in `config`.
13. **Upgrade the SDK before debugging a registration error.** "reuses a Registry name with a different identity" means the SDK is older than the gateway.
14. **Hosted judges need sessions.** App traffic is judgeable only after it is [recorded as a session](/evals/online#make-app-traffic-scoreable) whose `external_session_id` equals the spans' session ID.
15. **Compare structured values key-order-insensitively in TypeScript.** Stored `expected` values do not keep key order, so raw `JSON.stringify` equality scores `0`.
16. **Dataset names are workspace-wide.** Pushing to an existing name adds to that dataset; check `suite.created`.

## Task map

| Task | Page |
| - | - |
| First run, end to end | [Quickstart](/evals/quickstart) |
| Put cases in a versioned dataset | [Datasets](/evals/datasets) |
| Write a scorer: exact match, JSON, tolerance, pattern, LLM judge | [Code scorers](/evals/scorers) |
| Score tool calls, trajectories, multi-turn | [Evaluate an agent](/evals/agents) |
| Concurrency, async, resume, context manifest, reading results | [Run evals](/evals/offline) |
| Gate PRs, thresholds, verdicts, other CI systems | [Run evals in CI](/evals/ci) |
| Diff two versions case by case | [Compare two versions](/evals/cookbooks/compare-versions) |
| Keep an existing harness | [Bring your own runner](/evals/custom-runner) |
| Judge production sessions | [Hosted judges](/evals/judges), [score production sessions](/evals/online) |
| Log user feedback as scores | [Scores from your application](/evals/production-scores) |
| An error or a failed verdict | [Troubleshooting](/evals/troubleshooting) |
| Exact options, shapes, limits, routes | [Reference](/evals/reference) |

## Definition of done

Check each item. Each one can be verified by running a command or reading a file.

* [ ] `python -m evals.run` exits `0`, or `npx tsx evals/run.ts` does and `npx tsc --strict --noEmit` passes. Every case in `run.results` has no `error`.
* [ ] `run.summary["scores"]` (`run.summary.scores`) has one entry per scorer, each with a `count` equal to the number of cases the scorer applies to.
* [ ] `verify_experiment(client, run.experiment_id).raise_for_status()` passes, or in TypeScript `(await verifyExperiment(client, run.experimentId)).raiseForStatus()`: every result has a queryable trace, and the sampled traces carry their score spans.
* [ ] The eval name in `Eval(...)` is a constant, not derived from the branch or the time.
* [ ] No credential appears in the repository. `git grep -n "ATLAN_API_KEY="` finds only placeholders.
* [ ] `.github/workflows/atlan-evals.yml` triggers on `pull_request` and on `push` to the default branch, sets `ATLAN_API_KEY` and `ATLAN_WORKSPACE_ID` in the action step's `env`, grants only `contents: read` and `pull-requests: write`, and skips fork PRs.
* [ ] The human has been told to add the `ATLAN_API_KEY` secret and the `ATLAN_WORKSPACE_ID` variable, and to run the workflow once on `main` to create the baseline.
