> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Compare two agent versions

> Run a baseline and candidate over one dataset, then gate a release on the stored score summary.

Use the same dataset for both runs. Changing the dataset at the same time as the agent hides which change moved the score. This recipe uses a two-case synthetic suite; replace the task functions with calls to the two versions you want to compare.

## Create the cases once

Set `ATLAN_API_KEY` and `ATLAN_WORKSPACE_ID` as shown in [offline setup](/evals/offline#set-up), then run:

```python theme={null}
import os
from atlanai import AtlanClient, Eval, push_dataset

client = AtlanClient(
    os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
    bearer_token=os.environ["ATLAN_API_KEY"],
    workspace=os.environ["ATLAN_WORKSPACE_ID"],
)

suite = push_dataset(client, "release-regression-cases", [
    {"name": "arithmetic", "input": {"question": "What is 2 + 2?"}, "expected": {"value": "4"}},
    {"name": "geography", "input": {"question": "What is the capital of France?"}, "expected": {"value": "Paris"}},
])


def accuracy(*, output: str, expected: str, **_) -> float:
    return float(output == expected)


def baseline_task(question: str) -> str:
    return "4"  # Replace with the deployed version.


def candidate_task(question: str) -> str:
    return "4" if "2 + 2" in question else "Paris"  # Replace with the candidate.


baseline = Eval(
    "baseline", dataset=suite.dataset.id, task=baseline_task, scores=[accuracy],
    config={"agent_revision": "example-v1"},
)
candidate = Eval(
    "candidate", dataset=suite.dataset.id, task=candidate_task, scores=[accuracy],
    config={"agent_revision": "example-v2"},
    baseline_experiment_id=baseline.experiment_id,
)

base_score = baseline.summary["scores"]["accuracy"]["mean"]
new_score = candidate.summary["scores"]["accuracy"]["mean"]
print(baseline.experiment_id, base_score)
print(candidate.experiment_id, new_score)
if new_score < base_score:
    raise SystemExit("candidate regressed")
```

`baseline_experiment_id` records the relationship. The summary is calculated from all stored case results during finalization, so the gate reads the completed experiment rather than a partial page of results.

## Inspect a regression

If a real candidate scores lower, compare cases by `dataset_record_id`, then open the losing result's `trace_id`. Look for a changed prompt, model response, tool call, or retrieval result. Keep the scorer definition fixed during the comparison. If you change its rubric, create a new experiment and record which scorer version produced each score.

For nondeterministic tasks, run more than one trial per version. Store the trial and model settings in `config`; do not put each trial into a new dataset.
