> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# How evals work

> How evals measure whether an agent got better, what the registry stores, and how offline and online evals differ.

An eval answers one question about your agent with evidence you can open later:
did this change make it better or worse, and on which cases? Usage shows that
an agent ran. Evals show whether it did the job well.

## Three parts

Every eval has three parts:

* **Task**: your agent.
* **Data**: cases with inputs and, optionally, expected results.
* **Scores**: functions that turn one output into numbers.

The atlanai SDK runs every case in your own process or CI job, traces each run,
and stores the results as an **experiment**. The registry stores the evidence. It
never calls your agent.

## Offline and online evals

| | Offline evals | Online evals |
| - | - | - |
| Question | Did this change improve the agent on the cases I care about? | What is going wrong in real traffic? |
| Input | A fixed dataset or inline cases | Recorded sessions and their traces |
| Who runs it | Your process or CI job | A hosted judge in Agent Registry |
| Result | An experiment with one result per case | One insight per answer, linked to the session |

The two feed each other. Gate pull requests on an offline dataset, judge
production sessions to find failures it missed, then turn those failures into
new cases.

## What Agent Registry stores

* **Dataset**: a named set of cases in a workspace. Editing a case adds a new
  version.
* **Experiment**: one run over a dataset. Once it completes or fails, it is
  frozen.
* **Scorer**: what a score means. Every edit adds a version, and each score
  cites the exact version that produced it.
* **Trace**: the steps your agent took for each case.

Frozen experiments and versioned scorers keep runs comparable months apart.

## Example

A team changes its support agent's prompt and opens a pull request. CI runs the
agent over a 50-case dataset, scores each output, and stores an experiment.
The team compares its summary with the experiment from the main branch, and
opens the traces of the cases that got worse before merging.

## What evals do not do

* **Agent Registry does not run your agent.** Offline evals run in your process or CI
  job. There is no hosted runner.
* **Comparison runs in your tooling.** The registry stores each experiment's summary
  and results. The CI action and the compare cookbook compare two experiments
  case by case.
* **An experiment does not pin an agent version.** Record the version you
  tested in the experiment's configuration.

## See also

* [Run your first eval](/evals/quickstart): Score an agent against fixed cases and store an experiment.
* [Eval objects](/evals/concepts): Every eval object, what it holds, and when it is frozen.
* [Gate pull requests](/evals/ci): Run evals in CI and block changes that regress.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.