Skip to main content
An eval answers one question about your agent with evidence you can open later: did this change make it better or worse, and on which cases? Usage shows that an agent ran. Evals show whether it did the job well.

Three parts

Every eval has three parts:
  • Task: your agent.
  • Data: cases with inputs and, optionally, expected results.
  • Scores: functions that turn one output into numbers.
The atlanai SDK runs every case in your own process or CI job, traces each run, and stores the results as an experiment. The registry stores the evidence. It never calls your agent.

Offline and online evals

The two feed each other. Gate pull requests on an offline dataset, judge production sessions to find failures it missed, then turn those failures into new cases.

What Agent Registry stores

  • Dataset: a named set of cases in a workspace. Editing a case adds a new version.
  • Experiment: one run over a dataset. Once it completes or fails, it is frozen.
  • Scorer: what a score means. Every edit adds a version, and each score cites the exact version that produced it.
  • Trace: the steps your agent took for each case.
Frozen experiments and versioned scorers keep runs comparable months apart.

Example

A team changes its support agent’s prompt and opens a pull request. CI runs the agent over a 50-case dataset, scores each output, and stores an experiment. The team compares its summary with the experiment from the main branch, and opens the traces of the cases that got worse before merging.

What evals do not do

  • Agent Registry does not run your agent. Offline evals run in your process or CI job. There is no hosted runner.
  • Comparison runs in your tooling. The registry stores each experiment’s summary and results. The CI action and the compare cookbook compare two experiments case by case.
  • An experiment does not pin an agent version. Record the version you tested in the experiment’s configuration.

See also