Skip to main content
This page is written for a coding agent, such as Claude Code, Codex, or Cursor, that has been asked to set up evals for a project. It is equally a checklist for a person.

Load the docs

  • Every page is Markdown at <page>.md, for example https://platform.atlan.com/evals/quickstart.md. https://platform.atlan.com/llms.txt lists every page with its description.
  • Connect the docs MCP server so the agent can search current docs instead of relying on training data: claude mcp add --transport http atlan-docs https://platform.atlan.com/mcp. See For AI agents for other clients.
  • Read in this order: quickstart, concepts, scorers, agents, CI.

A brief to paste

Invariants

These are the facts an agent most often gets wrong. Each one is enforced by the API or the SDK.
  1. Atlan does not run the agent. Eval executes in the local process or CI job. There is no hosted runner to configure.
  2. Install the latest SDK. pip install -U 'atlanai[tracing]' for Python, npm install @atlanai/sdk@latest for TypeScript. The tracing extra is required in Python.
  3. Eval reads ATLAN_API_KEY, ATLAN_WORKSPACE_ID, and ATLAN_BASE_URL. AtlanClient reads nothing from the environment; pass its arguments explicitly.
  4. Keep the eval name fixed. Scorer identity and CI baselines key on it. Put variants in config.
  5. Scorers return 0..1, higher is better, or None to abstain. Never return a free 1.0 for a case a rule does not cover.
  6. Results are create-only and finished experiments are frozen. To fix a run, start a new one. summary.scores is computed by Atlan, not written by you.
  7. The task sees only input, and scorers see only what the task returns. Return the tool calls from the task if scorers need them. Record them at the tool, not from the agent’s self-report.
  8. Dataset records need an object input with the question under input_key (default question). {"value": x} as expected reaches the scorer as x.
  9. Hosted llm_judge scorers score recorded sessions, not Eval cases. For a judge inside Eval, call a model from a code scorer.
  10. Python Eval cannot run inside an event loop. Use await async_eval(...) in notebooks and async code.
  11. The CI action reads experiments that Python Eval reports. A TypeScript eval needs the results-file shim. The action’s api-key never reaches the command; set ATLAN_API_KEY and ATLAN_WORKSPACE_ID in the step’s env.
  12. An agent version is not pinned automatically. Record the commit and model in config.
  13. Upgrade the SDK before debugging a registration error. “reuses a Registry name with a different identity” means the SDK is older than the gateway.
  14. Hosted judges need sessions. App traffic is judgeable only after it is recorded as a session whose external_session_id equals the spans’ session ID.
  15. Compare structured values key-order-insensitively in TypeScript. Stored expected values do not keep key order, so raw JSON.stringify equality scores 0.
  16. Dataset names are workspace-wide. Pushing to an existing name adds to that dataset; check suite.created.

Task map

Definition of done

Check each item. Each one can be verified by running a command or reading a file.
  • python -m evals.run exits 0, or npx tsx evals/run.ts does and npx tsc --strict --noEmit passes. Every case in run.results has no error.
  • run.summary["scores"] (run.summary.scores) has one entry per scorer, each with a count equal to the number of cases the scorer applies to.
  • verify_experiment(client, run.experiment_id).raise_for_status() passes, or in TypeScript (await verifyExperiment(client, run.experimentId)).raiseForStatus(): every result has a queryable trace, and the sampled traces carry their score spans.
  • The eval name in Eval(...) is a constant, not derived from the branch or the time.
  • No credential appears in the repository. git grep -n "ATLAN_API_KEY=" finds only placeholders.
  • .github/workflows/atlan-evals.yml triggers on pull_request and on push to the default branch, sets ATLAN_API_KEY and ATLAN_WORKSPACE_ID in the action step’s env, grants only contents: read and pull-requests: write, and skips fork PRs.
  • The human has been told to add the ATLAN_API_KEY secret and the ATLAN_WORKSPACE_ID variable, and to run the workflow once on main to create the baseline.