> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Quickstart

> Run a two-case eval with the Python or TypeScript SDK, read its score summary, and open the trace behind a failed case.

This quickstart runs an eval with two cases and one scorer. One case is designed to fail, so you see what a failure looks like before you wire in a real agent. It takes about five minutes.

## 1. Install the SDK

Install the latest SDK. Python needs the `tracing` extra, because every eval case is traced. The TypeScript package includes tracing.

<CodeGroup>
  ```bash Python theme={null}
  pip install -U 'atlanai[tracing]'
  ```

  ```bash TypeScript theme={null}
  npm install @atlanai/sdk@latest
  ```
</CodeGroup>

## 2. Point it at your workspace

`Eval` reads three environment variables. Load the key from your secret store; do not paste it into source files or shell history.

```bash theme={null}
export ATLAN_API_KEY="<token-from-your-secret-store>"
export ATLAN_WORKSPACE_ID="workspace_01example"
# Only when you use a gateway other than https://api.atlan.com:
export ATLAN_BASE_URL="https://<gateway-host>"
```

The key must be able to create datasets, scorers, experiments, and results in that workspace. To find a workspace ID, run `atlanai workspace list`. [Authentication](/api/how-tos/authentication) explains which credential to use where.

## 3. Write the eval

An eval is three things. `data` is the cases. `task` is the function under test: it receives each case's `input` and returns an output. `scores` are the functions that grade each output. This task always answers `"4"`, so the second case fails.

<CodeGroup>
  ```python Python theme={null}
  # eval_quickstart.py
  from atlanai import Eval


  def answer(question: str) -> str:
      return "4"  # Replace with a call to your agent.


  def accuracy(*, output: str, expected: str, **_) -> float:
      return float(output == expected)


  run = Eval(
      "answer-regression",
      data=lambda: [
          {"id": "arithmetic", "input": "What is 2 + 2?", "expected": "4"},
          {"id": "geography", "input": "What is the capital of France?", "expected": "Paris"},
      ],
      task=answer,
      scores=[accuracy],
      config={"candidate": "example-v1"},
  )

  print(run.experiment_id, run.summary["scores"])
  for case in run.results:
      print(case.case_id, case.scores, case.trace_id)
  ```

  ```typescript TypeScript theme={null}
  // eval-quickstart.ts
  import { Eval } from "@atlanai/sdk";

  const run = await Eval("answer-regression", {
    data: () => [
      { id: "arithmetic", input: "What is 2 + 2?", expected: "4" },
      { id: "geography", input: "What is the capital of France?", expected: "Paris" },
    ],
    task: async (_question: string) => "4", // Replace with a call to your agent.
    scores: [{
      name: "accuracy",
      scorer: ({ output, expected }) => Number(output === expected),
    }],
  }, { config: { candidate: "example-v1" } });

  console.log(run.experimentId, run.summary.scores);
  for (const c of run.results) console.log(c.caseId, c.scores, c.traceId);
  ```
</CodeGroup>

## 4. Run it

<CodeGroup>
  ```bash Python theme={null}
  python eval_quickstart.py
  ```

  ```bash TypeScript theme={null}
  npx tsx eval-quickstart.ts
  ```
</CodeGroup>

You should see output like this. Your IDs will differ.

```text theme={null}
experiment_01example {'accuracy': {'mean': 0.5, 'count': 2}}
arithmetic {'accuracy': 1.0} 260d10c2be18fdf5f3ef6322a4d1430e
geography {'accuracy': 0.0} 2dff00b3bfd7cf25a2b0ce8409468d0c
```

Here is what happened:

1. `Eval` created an **experiment** with status `running`.
2. For each case it opened a trace, ran `answer`, and called `accuracy` on the output.
3. It checked that every trace had reached Atlan, then uploaded one **result** per case.
4. It marked the experiment `completed`. Atlan then computed `summary.scores`, the mean and count for each score name, from the stored results.

It also registered `accuracy` as a **scorer**, a versioned catalog entry. Each score records which scorer version produced it.

## 5. Look at the failure

The summary tells you *that* the run scored 0.5. The failed case's trace tells you *why*. Open the experiment in the Atlan app under **Evals → Experiments** and select the `geography` case. Or read the evidence back in code:

```python theme={null}
import os
from atlanai import AtlanClient

client = AtlanClient(
    os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
    bearer_token=os.environ["ATLAN_API_KEY"],
    workspace=os.environ["ATLAN_WORKSPACE_ID"],
)
results = client.experiments.results.list(run.experiment_id)
for r in results.items:
    print(r.case_id, r.output, r.scores, r.trace_id)
```

## Next steps

Replace `answer` with a call to your agent. Then work through these pages in order:

<CardGroup cols={2}>
  <Card title="Keep cases in a dataset" icon="table" href="/evals/datasets">
    Stable cases that every run and every PR is compared on.
  </Card>

  <Card title="Write better scorers" icon="list-check" href="/evals/scorers">
    Scorer signatures, return shapes, and a catalog of ready-to-copy scorers.
  </Card>

  <Card title="Evaluate an agent" icon="robot" href="/evals/agents">
    Score tool calls and trajectories, not just the final answer.
  </Card>

  <Card title="Run it in CI" icon="code-pull-request" href="/evals/ci">
    Post the score diff on every PR and fail on a regression.
  </Card>
</CardGroup>
