Skip to main content
Use the same dataset for both runs. Changing the dataset at the same time as the agent hides which change moved the score. This recipe uses a two-case synthetic suite; replace the task functions with calls to the two versions you want to compare.

Create the cases once

Set ATLAN_API_KEY and ATLAN_WORKSPACE_ID as shown in offline setup, then run:
baseline_experiment_id records the relationship. The summary is calculated from all stored case results during finalization, so the gate reads the completed experiment rather than a partial page of results.

Inspect a regression

If a real candidate scores lower, compare cases by dataset_record_id, then open the losing result’s trace_id. Look for a changed prompt, model response, tool call, or retrieval result. Keep the scorer definition fixed during the comparison. If you change its rubric, create a new experiment and record which scorer version produced each score. For nondeterministic tasks, run more than one trial per version. Store the trial and model settings in config; do not put each trial into a new dataset.