What you need
1. Write the eval script
The script builds the suite and runs it. It does not need to compare anything or choose an exit code: the action does that.evals/run.py
https://api.atlan.com, also set ATLAN_BASE_URL; the workflow does not need it, because the action works only with https://api.atlan.com.
Run it from the repository root as a module, python -m evals.run, with an empty evals/__init__.py, so that evals.scorers and your agent’s own package import. python evals/run.py puts evals/ itself on the import path, and from evals.scorers import ... then fails.
Inside the action, Eval stamps the commit, branch, and PR number onto the experiment config. It also reports each experiment to the action, so there is nothing else to wire up.
2. Add the workflow
.github/workflows/atlan-evals.yml
api-key input is used only by the action, to read results. It is deliberately not placed in the eval command’s environment, so the step’s env must set ATLAN_API_KEY and ATLAN_WORKSPACE_ID for Eval. Without them the command fails with a missing-credential error and the verdict is invalid. Your eval command runs in the same job, though, so treat it as trusted code. The action refuses to run under pull_request_target. @v0 follows reviewed 0.x releases. Pin a full commit SHA instead if your supply-chain policy requires immutable references.
The action reads from https://api.atlan.com only. To gate against another Atlan deployment, run your command in a plain step and use gate.py instead.
Do not add a
paths: filter to pull_request if the check is required by branch protection. A PR that touches no matching path never runs the workflow, so the required check stays pending and the PR cannot merge. To skip unrelated PRs, put the path test inside the job instead.3. Establish a baseline
Merge the workflow, or run it once onmain with Run workflow. That run becomes the baseline. Until one exists, PRs pass with a no-baseline warning.
The action picks the most recent completed run on the default branch with the same eval name (the experiment’s display_name) and dataset, and prefers one with the same model input. Set baseline to an experiment ID to pin a specific run, or to none to report the PR’s scores with no comparison.
What the PR comment shows
For each eval: a table of each score’s baseline mean, PR mean, and change, with counts of improved and regressed cases. It also shows informational cost, tokens, and latency, the verdict, and links to both experiments in Atlan. The comment lists record IDs and scores only. It never includes inputs, outputs, trace text, or error messages. It is updated in place on each push.Verdicts
Cost, tokens, and latency never affect the verdict.
Tune the gate
The action’s outputs,
eval-status, regression-count, experiment-ids, baseline-experiment-ids, and comment-id, let a later step act on the verdict.
Scores are assumed to be 0..1 with higher better. A score where lower is better, such as a step count, should be inverted in the scorer.
Run expensive suites on demand
Trigger only on a label, for a suite too slow or costly for every push:TypeScript evals
The action learns which experiments a command produced from a results file, and checks each experiment’s commit and PR. The Python SDK handles both automatically. A TypeScript eval does it in two lines: stamp the git identity intoconfig, and append the experiment to the results file.
evals/run.ts
display_name must be the eval name you pass to Eval: the action matches baselines on it. The action sets the ATLAN_EVAL_* variables for any command, so the shim works unchanged inside it.
The project needs tsx, typescript, and @types/node as dev dependencies (npm install -D tsx typescript @types/node), "type": "module" in package.json for top-level await, and a tsconfig.json such as:
tsconfig.json
.github/workflows/atlan-evals.yml
ATLAN_EVAL_RESULTS_FILE to a scratch path and ATLAN_EVAL_GIT_BRANCH/SHA/PR to test values, run the command, and read the line it appends.
Other CI systems
Outside GitHub Actions, gate with a short script. Run the eval, then compare it with a baseline experiment by the same rule the action uses, and exit non-zero on a regression.evals/gate.py
Eval does not record the branch, so pass it yourself, for example config={"branch": os.environ["CI_BRANCH"]}, and filter on it: client.experiments.list(config="branch:main", dataset_id=suite.id, experiment_status="completed", sort="-created_at", limit=1). config filters take a dotted path; runs made by the action carry the branch at git.branch. The compare cookbook covers per-case diffs in more depth.