> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Run evals in CI

> Run your eval on every pull request, compare it with the default branch, post the score diff as a PR comment, and fail the check on a regression.

A CI eval answers one question on every pull request: *did this change make the agent worse on the cases we care about?* The Atlan Evals GitHub Action runs your eval command and finds the last run on the default branch. It compares the two case by case, comments the diff on the PR, and fails the check on a regression.

```text theme={null}
push to main ─► eval run ─► becomes the baseline
pull request ─► eval run ─► compared with the baseline ─► PR comment + pass/fail check
```

## What you need

| | Detail |
| - | - |
| An eval script | A command such as `python -m evals.run` that runs one or more `Eval` calls. Keep each eval's name the same on every branch: baselines are matched on the name and the dataset. |
| A dataset | Recommended. Results are paired across runs by `dataset_record_id`, or by `case_id` for inline cases with stable IDs. |
| An Atlan API key | Workspace-scoped, able to create experiments in the workspace. Store it as a repository secret. |
| The workspace ID | Store it as a repository variable. |

## 1. Write the eval script

The script builds the suite and runs it. It does not need to compare anything or choose an exit code: the action does that.

```python evals/run.py theme={null}
import os
from atlanai import AtlanClient, Eval, push_dataset

from my_agent import answer               # your agent
from evals.scorers import exact_match, required_facts

client = AtlanClient(
    os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
    bearer_token=os.environ["ATLAN_API_KEY"],
    workspace=os.environ["ATLAN_WORKSPACE_ID"],
)

suite = push_dataset(client, "support-agent-cases", [
    {"name": "refund-status", "input": {"question": "Where is my refund for order 1042?"},
     "expected": {"value": "processed"}},
    # ... the rest of your cases, or load them from a file in the repo
])

Eval(
    "support-agent",                    # the same name on every branch
    dataset=suite.id,
    task=answer,
    scores=[exact_match, required_facts],
    client=client,
)
```

To run the same command locally against a gateway other than `https://api.atlan.com`, also set `ATLAN_BASE_URL`; the workflow does not need it, because the action works only with `https://api.atlan.com`.

Run it from the repository root as a module, `python -m evals.run`, with an empty `evals/__init__.py`, so that `evals.scorers` and your agent's own package import. `python evals/run.py` puts `evals/` itself on the import path, and `from evals.scorers import ...` then fails.

Inside the action, `Eval` stamps the commit, branch, and PR number onto the experiment config. It also reports each experiment to the action, so there is nothing else to wire up.

## 2. Add the workflow

```yaml .github/workflows/atlan-evals.yml theme={null}
name: Atlan evals

on:
  pull_request:
  push:
    branches: [main]            # each run on main becomes the next baseline
  workflow_dispatch:

permissions:
  contents: read
  pull-requests: write          # to post the comment

concurrency:
  group: atlan-evals-${{ github.event.pull_request.number || github.ref }}
  cancel-in-progress: true

jobs:
  eval:
    # Fork PRs never receive secrets; skip them rather than fail confusingly.
    if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
        with:
          persist-credentials: false
      - uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5.6.0
        with:
          python-version: "3.12"
      - run: pip install -U 'atlanai[tracing]' -r requirements.txt   # your agent's dependencies
      - uses: atlanai/agent-registry-action/eval@v0
        with:
          command: python -m evals.run
          api-key: ${{ secrets.ATLAN_API_KEY }}
          workspace-id: ${{ vars.ATLAN_WORKSPACE_ID }}
          thresholds: |
            default=0.05
        env:
          # The command needs its own credentials: the action does not pass api-key to it.
          ATLAN_API_KEY: ${{ secrets.ATLAN_API_KEY }}
          ATLAN_WORKSPACE_ID: ${{ vars.ATLAN_WORKSPACE_ID }}
          # Secrets your agent needs go here too.
          MODEL_API_KEY: ${{ secrets.MODEL_API_KEY }}
```

The `api-key` input is used only by the action, to read results. It is deliberately not placed in the eval command's environment, so the step's `env` must set `ATLAN_API_KEY` and `ATLAN_WORKSPACE_ID` for `Eval`. Without them the command fails with a missing-credential error and the verdict is `invalid`. Your eval command runs in the same job, though, so treat it as trusted code. The action refuses to run under `pull_request_target`. `@v0` follows reviewed `0.x` releases. Pin a full commit SHA instead if your supply-chain policy requires immutable references.

The action reads from `https://api.atlan.com` only. To gate against another Atlan deployment, run your command in a plain step and use [`gate.py`](#other-ci-systems) instead.

<Note>
  Do not add a `paths:` filter to `pull_request` if the check is required by branch protection. A PR that touches no matching path never runs the workflow, so the required check stays pending and the PR cannot merge. To skip unrelated PRs, put the path test inside the job instead.
</Note>

## 3. Establish a baseline

Merge the workflow, or run it once on `main` with **Run workflow**. That run becomes the baseline. Until one exists, PRs pass with a `no-baseline` warning.

The action picks the most recent `completed` run on the default branch with the same eval name (the experiment's `display_name`) and dataset, and prefers one with the same `model` input. Set `baseline` to an experiment ID to pin a specific run, or to `none` to report the PR's scores with no comparison.

## What the PR comment shows

For each eval: a table of each score's baseline mean, PR mean, and change, with counts of improved and regressed cases. It also shows informational cost, tokens, and latency, the verdict, and links to both experiments in Atlan. The comment lists record IDs and scores only. It never includes inputs, outputs, trace text, or error messages. It is updated in place on each push.

## Verdicts

| Verdict | When | Check |
| - | - | - |
| `invalid` | The command exited non-zero or reported no experiment. Or the run had no cases or no scores, more than `max-error-rate` of its cases errored, it did not complete, it came from another commit or PR, or its results could not be read in full. | Fails, always |
| `regression` | A score's mean dropped by more than its threshold, **and** at least `min-regressed-cases` cases got worse | Fails, unless `fail-on: invalid` |
| `no-baseline` | No baseline, or no cases shared with it | Passes, with a warning |
| `pass` | Otherwise | Passes |

Cost, tokens, and latency never affect the verdict.

## Tune the gate

| Input | Default | Use |
| - | - | - |
| `thresholds` | `default=0.05` | How far a mean may drop, one line per score: `default=0.05`, `exact_match=0.02`, `tone=off` |
| `min-regressed-cases` | `2` | Guards against one flaky case failing the build |
| `max-error-rate` | `0.2` | Share of cases that may error before the run counts as broken |
| `fail-on` | `regression` | `invalid` reports regressions without failing on them, which is useful while you calibrate |
| `baseline` | `auto` | `auto`, `none`, or an experiment ID |
| `baseline-branch` | the default branch | Compare against a release branch instead |
| `model` | none | Record and prefer baselines for a specific model |

The action's outputs, `eval-status`, `regression-count`, `experiment-ids`, `baseline-experiment-ids`, and `comment-id`, let a later step act on the verdict.

Scores are assumed to be `0..1` with higher better. A score where lower is better, such as a step count, should be inverted in the scorer.

## Run expensive suites on demand

Trigger only on a label, for a suite too slow or costly for every push:

```yaml theme={null}
on:
  pull_request:
    types: [labeled, synchronize]
jobs:
  eval:
    if: contains(github.event.pull_request.labels.*.name, 'run-evals')
```

## TypeScript evals

The action learns which experiments a command produced from a results file, and checks each experiment's commit and PR. The Python SDK handles both automatically. A TypeScript eval does it in two lines: stamp the git identity into `config`, and append the experiment to the results file.

```typescript evals/run.ts theme={null}
import { appendFileSync } from "node:fs";
import { Eval } from "@atlanai/sdk";

const env = process.env;
const git = Object.fromEntries(
  [["branch", env.ATLAN_EVAL_GIT_BRANCH], ["sha", env.ATLAN_EVAL_GIT_SHA], ["pr", env.ATLAN_EVAL_GIT_PR]]
    .filter(([, v]) => v),
);

const run = await Eval("support-agent", { dataset: "support-agent-cases", task: answer, scores: [exactMatch] }, {
  config: { git, ...(env.ATLAN_EVAL_MODEL ? { model: env.ATLAN_EVAL_MODEL } : {}) },
});

if (env.ATLAN_EVAL_RESULTS_FILE) {
  const e = run.experiment as { id: string; name: string; displayName?: string; datasetId?: string };
  appendFileSync(env.ATLAN_EVAL_RESULTS_FILE, JSON.stringify({
    experiment_id: run.experimentId, name: e.name, display_name: e.displayName ?? "support-agent",
    dataset_id: e.datasetId ?? null, status: "completed",
  }) + "\n");
}
```

`display_name` must be the eval name you pass to `Eval`: the action matches baselines on it. The action sets the `ATLAN_EVAL_*` variables for any command, so the shim works unchanged inside it.

The project needs `tsx`, `typescript`, and `@types/node` as dev dependencies (`npm install -D tsx typescript @types/node`), `"type": "module"` in `package.json` for top-level `await`, and a `tsconfig.json` such as:

```json tsconfig.json theme={null}
{
  "compilerOptions": {
    "target": "es2022",
    "module": "nodenext",
    "moduleResolution": "nodenext",
    "strict": true,
    "noEmit": true,
    "skipLibCheck": true
  },
  "include": ["src", "evals"]
}
```

The complete workflow:

```yaml .github/workflows/atlan-evals.yml theme={null}
name: Atlan evals

on:
  pull_request:
  push:
    branches: [main]
  workflow_dispatch:

permissions:
  contents: read
  pull-requests: write

concurrency:
  group: atlan-evals-${{ github.event.pull_request.number || github.ref }}
  cancel-in-progress: true

jobs:
  eval:
    if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
        with:
          persist-credentials: false
      - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
        with:
          node-version: "22"
      - run: npm ci
      - run: npx tsc --noEmit
      - uses: atlanai/agent-registry-action/eval@v0
        with:
          command: npx tsx evals/run.ts
          api-key: ${{ secrets.ATLAN_API_KEY }}
          workspace-id: ${{ vars.ATLAN_WORKSPACE_ID }}
          thresholds: |
            default=0.05
        env:
          ATLAN_API_KEY: ${{ secrets.ATLAN_API_KEY }}
          ATLAN_WORKSPACE_ID: ${{ vars.ATLAN_WORKSPACE_ID }}
          MODEL_API_KEY: ${{ secrets.MODEL_API_KEY }}
```

To check the shim locally, set `ATLAN_EVAL_RESULTS_FILE` to a scratch path and `ATLAN_EVAL_GIT_BRANCH`/`SHA`/`PR` to test values, run the command, and read the line it appends.

## Other CI systems

Outside GitHub Actions, gate with a short script. Run the eval, then compare it with a baseline experiment by the same rule the action uses, and exit non-zero on a regression.

```python evals/gate.py theme={null}
"""Fail the build when a candidate experiment regresses against a baseline.

    python evals/gate.py <baseline_experiment_id> <candidate_experiment_id>

Exit 0 on pass, 1 on a regression, 2 on a broken run.
"""
import os
import sys
from atlanai import AtlanClient

THRESHOLD = 0.05          # a score's mean may drop by at most this much
MIN_REGRESSED_CASES = 2   # ...and at least this many cases must get worse
MAX_ERROR_RATE = 0.2

client = AtlanClient(
    os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
    bearer_token=os.environ["ATLAN_API_KEY"],
    workspace=os.environ["ATLAN_WORKSPACE_ID"],
)


def results(experiment_id):
    rows, offset = [], 0
    while True:
        page = client.experiments.results.list(experiment_id, limit=500, offset=offset).items
        rows += page
        if len(page) < 500:
            return {r.dataset_record_id or r.case_id: r for r in rows}
        offset += 500


baseline_id, candidate_id = sys.argv[1], sys.argv[2]
if client.experiments.get(candidate_id).experiment_status != "completed":
    print("broken run: the candidate did not complete")
    sys.exit(2)
base, cand = results(baseline_id), results(candidate_id)
errors = sum(1 for r in cand.values() if r.error)
if not cand or errors / len(cand) > MAX_ERROR_RATE:
    print(f"broken run: {errors}/{len(cand)} cases errored")
    sys.exit(2)

shared = base.keys() & cand.keys()
names = {n for r in cand.values() for n in (r.scores or {})}
regressed = False
for name in sorted(names):
    pairs = [((base[k].scores or {}).get(name), (cand[k].scores or {}).get(name)) for k in shared]
    pairs = [(b, c) for b, c in pairs if isinstance(b, (int, float)) and isinstance(c, (int, float))]
    if not pairs:
        continue
    b_mean = sum(b for b, _ in pairs) / len(pairs)
    c_mean = sum(c for _, c in pairs) / len(pairs)
    worse = sum(1 for b, c in pairs if c < b)
    bad = b_mean - c_mean > THRESHOLD and worse >= MIN_REGRESSED_CASES
    regressed |= bad
    print(f"{'REGRESSED' if bad else 'ok':9} {name}: {b_mean:.3f} -> {c_mean:.3f} ({worse} of {len(pairs)} cases worse)")
sys.exit(1 if regressed else 0)
```

```text theme={null}
REGRESSED exact_match: 1.000 -> 0.500 (3 of 6 cases worse)
```

To find the baseline, list the latest completed run of the same eval and dataset on your main branch. Outside the GitHub Action, `Eval` does not record the branch, so pass it yourself, for example `config={"branch": os.environ["CI_BRANCH"]}`, and filter on it: `client.experiments.list(config="branch:main", dataset_id=suite.id, experiment_status="completed", sort="-created_at", limit=1)`. `config` filters take a dotted path; runs made by the action carry the branch at `git.branch`. The [compare cookbook](/evals/cookbooks/compare-versions) covers per-case diffs in more depth.
