1. Install the SDK
Install the latest SDK. Python needs thetracing extra, because every eval case is traced. The TypeScript package includes tracing.
2. Point it at your workspace
Eval reads three environment variables. Load the key from your secret store; do not paste it into source files or shell history.
atlanai workspace list. Authentication explains which credential to use where.
3. Write the eval
An eval is three things.data is the cases. task is the function under test: it receives each case’s input and returns an output. scores are the functions that grade each output. This task always answers "4", so the second case fails.
4. Run it
Evalcreated an experiment with statusrunning.- For each case it opened a trace, ran
answer, and calledaccuracyon the output. - It checked that every trace had reached Atlan, then uploaded one result per case.
- It marked the experiment
completed. Atlan then computedsummary.scores, the mean and count for each score name, from the stored results.
accuracy as a scorer, a versioned catalog entry. Each score records which scorer version produced it.
5. Look at the failure
The summary tells you that the run scored 0.5. The failed case’s trace tells you why. Open the experiment in the Atlan app under Evals → Experiments and select thegeography case. Or read the evidence back in code:
Next steps
Replaceanswer with a call to your agent. Then work through these pages in order:
Keep cases in a dataset
Stable cases that every run and every PR is compared on.
Write better scorers
Scorer signatures, return shapes, and a catalog of ready-to-copy scorers.
Evaluate an agent
Score tool calls and trajectories, not just the final answer.
Run it in CI
Post the score diff on every PR and fail on a regression.