Skip to main content
Every eval object answers one question. Knowing which objects are versioned, create-only, or frozen tells you what you can rely on when you compare runs months apart.

The objects

Snapshots make runs comparable

When an experiment starts from a dataset, Atlan stamps a dataset_snapshot onto it: the dataset version plus each live record’s ID, version, content hash, input, and expected value. The run executes from that snapshot. Editing a record later changes only future experiments, so two runs on the same dataset compare the same cases unless you deliberately changed them. Results join across experiments on dataset_record_id. That is the key for “which cases got better or worse”.

Where the summary comes from

summary.scores is computed by Atlan, not by your runner. When the experiment moves from running to completed or failed, Atlan reads every stored result and replaces summary.scores with {name: {mean, count}} for each score name. A null score, meaning “not applicable”, is left out of both the mean and the count. Other summary fields your runner sends, such as case counts and duration, are kept as sent. Use the stored summary for release decisions and run history. Do not average a page of results yourself: a partial page gives a partial mean.

Offline and online scores are stored differently

A score span you write from application code, for example a user’s thumbs-up, is a third path. See scores from your application.

What Atlan does and does not do

  • Atlan stores and computes. It stores datasets, snapshots, results, scorer versions, and traces. It computes summaries. It executes hosted llm_judge scorers over recorded sessions.
  • Atlan does not run your agent. Offline evals execute in your process or CI job. There is no hosted runner, and no Atlan service calls your model or tools.
  • There is no compare endpoint. A baseline is a stored pointer. You compare two experiments by reading their summaries and joining their results on dataset_record_id. The CI action and the compare cookbook both do this for you.
  • An experiment does not pin an agent version. Record the version you tested in the experiment config, or pin each input with a context manifest.

Names you will see in code