> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Evals

> Test an agent before release and score the sessions it produces in use.

An eval asks a concrete question about an agent run and keeps the evidence behind the answer. Atlan supports two ways to ask it:

| | Offline evals | Online evals |
| - | - | - |
| Input | A fixed set of cases, with optional expected answers | Recorded sessions and their traces |
| When | During development or CI | After real sessions complete |
| Execution | Your code runs the task and scorers with `Eval` | Agent Gateway runs a stored `llm_judge` scorer |
| Durable result | Experiment, case results, score summary, and traces | One `session_insight` per answer, linked to the session and scorer version |
| Best first question | Did the candidate improve on the same cases? | What is failing in real traffic? |

Start with a small offline dataset that represents the behavior you need. Compare candidates against the same cases. Then score completed sessions to find failure modes the dataset missed. Add reviewed failures back to the dataset.

```text theme={null}
curated cases → offline experiment → compare and inspect traces
                                          ↑
completed sessions → online scores → review a failure
```

<CardGroup cols={2}>
  <Card title="Run an offline eval" icon="flask" href="/evals/offline">
    Install the tracing SDK, run cases, read the summary, and resume an interrupted run.
  </Card>

  <Card title="Score production sessions" icon="chart-line" href="/evals/online">
    Define a judge, preview its prompt, test it, and start an asynchronous scoring run.
  </Card>

  <Card title="Compare two agent versions" icon="code-branch" href="/evals/cookbooks/compare-versions">
    Use one dataset for a baseline and a candidate so each case has a direct comparison.
  </Card>

  <Card title="Turn a failure into a test" icon="rotate" href="/evals/cookbooks/production-to-dataset">
    Review a low score and add a reproducible case to the next offline run.
  </Card>
</CardGroup>

## The objects you will see

* A **dataset** holds curated cases. Each record keeps its input, optional expected result, categories, and provenance.
* An **experiment** is one offline run over a dataset or inline cases. It pins the run configuration and records the case results.
* A **scorer** defines what a score means. Offline `Eval` calls your score function. Online scoring runs a stored `llm_judge` definition against session evidence.
* A **trace** shows the task, model, and tool work behind a result. An offline result points to its case trace; an online insight points to the judged session or trace.

Keep prompts, tool arguments, and session content within your approved data policy. A hosted online judge receives the selected session transcript. [Tracing privacy controls](/tools/sdk/how-tos/privacy) determine what your application sends to Atlan in the first place.
