The objects
Snapshots make runs comparable
When an experiment starts from a dataset, Atlan stamps adataset_snapshot onto it: the dataset version plus each live record’s ID, version, content hash, input, and expected value. The run executes from that snapshot. Editing a record later changes only future experiments, so two runs on the same dataset compare the same cases unless you deliberately changed them.
Results join across experiments on dataset_record_id. That is the key for “which cases got better or worse”.
Where the summary comes from
summary.scores is computed by Atlan, not by your runner. When the experiment moves from running to completed or failed, Atlan reads every stored result and replaces summary.scores with {name: {mean, count}} for each score name. A null score, meaning “not applicable”, is left out of both the mean and the count. Other summary fields your runner sends, such as case counts and duration, are kept as sent.
Use the stored summary for release decisions and run history. Do not average a page of results yourself: a partial page gives a partial mean.
Offline and online scores are stored differently
A score span you write from application code, for example a user’s thumbs-up, is a third path. See scores from your application.
What Atlan does and does not do
- Atlan stores and computes. It stores datasets, snapshots, results, scorer versions, and traces. It computes summaries. It executes hosted
llm_judgescorers over recorded sessions. - Atlan does not run your agent. Offline evals execute in your process or CI job. There is no hosted runner, and no Atlan service calls your model or tools.
- There is no compare endpoint. A baseline is a stored pointer. You compare two experiments by reading their summaries and joining their results on
dataset_record_id. The CI action and the compare cookbook both do this for you. - An experiment does not pin an agent version. Record the version you tested in the experiment
config, or pin each input with a context manifest.