Skip to main content
An agent can reach the right answer by the wrong path, or the wrong answer by a reasonable one. Score both: the final output, and the trajectory, meaning which tools it called, with which arguments, in how many steps. The pattern has three parts:
  1. The task returns the trajectory, not just the answer. Your scorers see only what the task returns, so return the tool calls alongside the answer.
  2. The trace records everything else. Model calls, tool spans, tokens, cost, and latency land on the case’s trace for diagnosis, without you returning them.
  3. The dataset states what a good trajectory is. Put the expected tools and arguments in expected.

Wrap the agent as a task

Give your agent a thin adapter that returns a structured output: the answer, plus the tool calls the agent made. Record the tool calls at the tool boundary, where they execute, not from what the agent says it did. A model’s own account of its calls can differ from what ran. Here each tool is traced and records its own calls into a per-case list. The list is a context variable, so concurrent cases do not mix their calls.
project_name is the service name on your traces; there is nothing to create in Atlan first. Create the logger yourself when you instrument your agent, and pass it to Eval as trace_logger in Python or logger in TypeScript (see run it). Its spans then nest under each case. Every case trace looks like this:
Model and framework spans appear when their instrumentation is installed. In Python, auto_instrument() enables installed OpenTelemetry instrumentors for providers such as OpenAI and Anthropic, and frameworks such as LangChain, CrewAI, and the OpenAI Agents SDK. In TypeScript, wrap the Vercel AI SDK with wrapAISDK. See tracing integrations for the full list. Every eval case is traced in full: Eval refuses a logger that samples or drops content.

Stub tools with side effects

An eval runs your agent for real. A tool that sends email, issues a refund, or writes to production does that on every case of every run. Give the agent a stand-in for each such tool: it records the call, so the scorers still see it, and returns a plausible result without acting.
Pass the stand-ins where your agent looks up its tools, for example its tool registry or constructor. Keep each stand-in’s name and signature identical to the real tool’s, so the tool schema the model sees does not change.

Score tool calls

These scorers read output["tool_calls"] and expected["tools"] as lists of {name, args}; in TypeScript, output.toolCalls. Adapt the field names to your adapter.
In TypeScript, declare expected optional and return null when it is missing: Eval types it as possibly absent, and a required expected fails tsc --strict. In TypeScript, compare structured values with same, not raw JSON.stringify: expected read from a dataset comes back with its object keys in a different order, and a raw comparison then scores 0 on every case. tool_arguments matches each expected call to at most one actual call, so an agent that calls a tool once cannot satisfy two expected calls to it. Which to use: Compare arguments on the fields that matter. Normalise or drop volatile ones, such as timestamps and request IDs, before comparing, or the score measures noise.

Run it

Put the agent, the scorers, and this call in one file and run it.

Evaluate a multi-turn conversation

Put the user turns in input as a list. The task replays them against your agent, passing the history so far, and returns the transcript and every tool call across the turns. It returns the same answer and tool_calls shape as the single-turn task, with the final reply as answer, so the scorers above work on both. Scorers can then judge the final turn, every turn, or the whole trajectory. To serve single-turn and multi-turn records from one dataset, normalise first: turns = question if isinstance(question, list) else [question].
In a dataset, store the turns under the input_key, for example {"question": ["...", "..."]}, so the task receives the list.

Tie the run to the agent

Set the subject so the run appears on that agent’s profile in the Atlan app, and so you can list its history:
subject_kind is agent, harness, or skill, and subject_id must be an artifact of that kind you can read. The experiment does not record which version of the agent ran. Put the version, commit, model, and prompt revision in config, or pin them with a context manifest.

Cost, tokens, and latency

Cost and tokens live on the traces, not in the summary. Query them per experiment with trace stats. A scalar query returns one number per measure; a table query grouped by score_name returns the mean of each score across the run’s score spans.
Token and cost values come from model spans that report usage. An agent without instrumented model calls shows 0. To enforce a budget, compare these numbers in your CI script. The CI action reports them for information and never fails a check on them.

Nondeterminism

Agents vary run to run. Eval runs each case once. To measure variance, run the same eval several times with a trial field in config, and compare the means:
Set model temperature to 0 where your agent allows it. A gate threshold narrower than the run-to-run spread fails at random. Widen it, or gate on more cases.