Load the docs
- Every page is Markdown at
<page>.md, for examplehttps://platform.atlan.com/evals/quickstart.md.https://platform.atlan.com/llms.txtlists every page with its description. - Connect the docs MCP server so the agent can search current docs instead of relying on training data:
claude mcp add --transport http atlan-docs https://platform.atlan.com/mcp. See For AI agents for other clients. - Read in this order: quickstart, concepts, scorers, agents, CI.
A brief to paste
Invariants
These are the facts an agent most often gets wrong. Each one is enforced by the API or the SDK.- Atlan does not run the agent.
Evalexecutes in the local process or CI job. There is no hosted runner to configure. - Install the latest SDK.
pip install -U 'atlanai[tracing]'for Python,npm install @atlanai/sdk@latestfor TypeScript. Thetracingextra is required in Python. EvalreadsATLAN_API_KEY,ATLAN_WORKSPACE_ID, andATLAN_BASE_URL.AtlanClientreads nothing from the environment; pass its arguments explicitly.- Keep the eval name fixed. Scorer identity and CI baselines key on it. Put variants in
config. - Scorers return
0..1, higher is better, orNoneto abstain. Never return a free1.0for a case a rule does not cover. - Results are create-only and finished experiments are frozen. To fix a run, start a new one.
summary.scoresis computed by Atlan, not written by you. - The task sees only
input, and scorers see only what the task returns. Return the tool calls from the task if scorers need them. Record them at the tool, not from the agent’s self-report. - Dataset records need an object
inputwith the question underinput_key(defaultquestion).{"value": x}asexpectedreaches the scorer asx. - Hosted
llm_judgescorers score recorded sessions, notEvalcases. For a judge insideEval, call a model from a code scorer. - Python
Evalcannot run inside an event loop. Useawait async_eval(...)in notebooks and async code. - The CI action reads experiments that Python
Evalreports. A TypeScript eval needs the results-file shim. The action’sapi-keynever reaches the command; setATLAN_API_KEYandATLAN_WORKSPACE_IDin the step’senv. - An agent version is not pinned automatically. Record the commit and model in
config. - Upgrade the SDK before debugging a registration error. “reuses a Registry name with a different identity” means the SDK is older than the gateway.
- Hosted judges need sessions. App traffic is judgeable only after it is recorded as a session whose
external_session_idequals the spans’ session ID. - Compare structured values key-order-insensitively in TypeScript. Stored
expectedvalues do not keep key order, so rawJSON.stringifyequality scores0. - Dataset names are workspace-wide. Pushing to an existing name adds to that dataset; check
suite.created.
Task map
Definition of done
Check each item. Each one can be verified by running a command or reading a file.-
python -m evals.runexits0, ornpx tsx evals/run.tsdoes andnpx tsc --strict --noEmitpasses. Every case inrun.resultshas noerror. -
run.summary["scores"](run.summary.scores) has one entry per scorer, each with acountequal to the number of cases the scorer applies to. -
verify_experiment(client, run.experiment_id).raise_for_status()passes, or in TypeScript(await verifyExperiment(client, run.experimentId)).raiseForStatus(): every result has a queryable trace, and the sampled traces carry their score spans. - The eval name in
Eval(...)is a constant, not derived from the branch or the time. - No credential appears in the repository.
git grep -n "ATLAN_API_KEY="finds only placeholders. -
.github/workflows/atlan-evals.ymltriggers onpull_requestand onpushto the default branch, setsATLAN_API_KEYandATLAN_WORKSPACE_IDin the action step’senv, grants onlycontents: readandpull-requests: write, and skips fork PRs. - The human has been told to add the
ATLAN_API_KEYsecret and theATLAN_WORKSPACE_IDvariable, and to run the workflow once onmainto create the baseline.