Three parts
Every eval has three parts:- Task: your agent.
- Data: cases with inputs and, optionally, expected results.
- Scores: functions that turn one output into numbers.
Offline and online evals
The two feed each other. Gate pull requests on an offline dataset, judge
production sessions to find failures it missed, then turn those failures into
new cases.
What Agent Registry stores
- Dataset: a named set of cases in a workspace. Editing a case adds a new version.
- Experiment: one run over a dataset. Once it completes or fails, it is frozen.
- Scorer: what a score means. Every edit adds a version, and each score cites the exact version that produced it.
- Trace: the steps your agent took for each case.
Example
A team changes its support agent’s prompt and opens a pull request. CI runs the agent over a 50-case dataset, scores each output, and stores an experiment. The team compares its summary with the experiment from the main branch, and opens the traces of the cases that got worse before merging.What evals do not do
- Agent Registry does not run your agent. Offline evals run in your process or CI job. There is no hosted runner.
- Comparison runs in your tooling. The registry stores each experiment’s summary and results. The CI action and the compare cookbook compare two experiments case by case.
- An experiment does not pin an agent version. Record the version you tested in the experiment’s configuration.
See also
- Run your first eval: Score an agent against fixed cases and store an experiment.
- Eval objects: Every eval object, what it holds, and when it is frozen.
- Gate pull requests: Run evals in CI and block changes that regress.