Skip to main content
This quickstart runs an eval with two cases and one scorer. One case is designed to fail, so you see what a failure looks like before you wire in a real agent. It takes about five minutes.

1. Install the SDK

Install the latest SDK. Python needs the tracing extra, because every eval case is traced. The TypeScript package includes tracing.

2. Point it at your workspace

Eval reads three environment variables. Load the key from your secret store; do not paste it into source files or shell history.
The key must be able to create datasets, scorers, experiments, and results in that workspace. To find a workspace ID, run atlanai workspace list. Authentication explains which credential to use where.

3. Write the eval

An eval is three things. data is the cases. task is the function under test: it receives each case’s input and returns an output. scores are the functions that grade each output. This task always answers "4", so the second case fails.

4. Run it

You should see output like this. Your IDs will differ.
Here is what happened:
  1. Eval created an experiment with status running.
  2. For each case it opened a trace, ran answer, and called accuracy on the output.
  3. It checked that every trace had reached Atlan, then uploaded one result per case.
  4. It marked the experiment completed. Atlan then computed summary.scores, the mean and count for each score name, from the stored results.
It also registered accuracy as a scorer, a versioned catalog entry. Each score records which scorer version produced it.

5. Look at the failure

The summary tells you that the run scored 0.5. The failed case’s trace tells you why. Open the experiment in the Atlan app under Evals → Experiments and select the geography case. Or read the evidence back in code:

Next steps

Replace answer with a call to your agent. Then work through these pages in order:

Keep cases in a dataset

Stable cases that every run and every PR is compared on.

Write better scorers

Scorer signatures, return shapes, and a catalog of ready-to-copy scorers.

Evaluate an agent

Score tool calls and trajectories, not just the final answer.

Run it in CI

Post the score diff on every PR and fail on a regression.