- The task returns the trajectory, not just the answer. Your scorers see only what the task returns, so return the tool calls alongside the answer.
- The trace records everything else. Model calls, tool spans, tokens, cost, and latency land on the case’s trace for diagnosis, without you returning them.
- The dataset states what a good trajectory is. Put the expected tools and arguments in
expected.
Wrap the agent as a task
Give your agent a thin adapter that returns a structured output: the answer, plus the tool calls the agent made. Record the tool calls at the tool boundary, where they execute, not from what the agent says it did. A model’s own account of its calls can differ from what ran. Here each tool is traced and records its own calls into a per-case list. The list is a context variable, so concurrent cases do not mix their calls.project_name is the service name on your traces; there is nothing to create in Atlan first. Create the logger yourself when you instrument your agent, and pass it to Eval as trace_logger in Python or logger in TypeScript (see run it). Its spans then nest under each case. Every case trace looks like this:
auto_instrument() enables installed OpenTelemetry instrumentors for providers such as OpenAI and Anthropic, and frameworks such as LangChain, CrewAI, and the OpenAI Agents SDK. In TypeScript, wrap the Vercel AI SDK with wrapAISDK. See tracing integrations for the full list. Every eval case is traced in full: Eval refuses a logger that samples or drops content.
Stub tools with side effects
An eval runs your agent for real. A tool that sends email, issues a refund, or writes to production does that on every case of every run. Give the agent a stand-in for each such tool: it records the call, so the scorers still see it, and returns a plausible result without acting.Score tool calls
These scorers readoutput["tool_calls"] and expected["tools"] as lists of {name, args}; in TypeScript, output.toolCalls. Adapt the field names to your adapter.
expected optional and return null when it is missing: Eval types it as possibly absent, and a required expected fails tsc --strict. In TypeScript, compare structured values with same, not raw JSON.stringify: expected read from a dataset comes back with its object keys in a different order, and a raw comparison then scores 0 on every case. tool_arguments matches each expected call to at most one actual call, so an agent that calls a tool once cannot satisfy two expected calls to it.
Which to use:
Compare arguments on the fields that matter. Normalise or drop volatile ones, such as timestamps and request IDs, before comparing, or the score measures noise.
Run it
Put the agent, the scorers, and this call in one file and run it.Evaluate a multi-turn conversation
Put the user turns ininput as a list. The task replays them against your agent, passing the history so far, and returns the transcript and every tool call across the turns. It returns the same answer and tool_calls shape as the single-turn task, with the final reply as answer, so the scorers above work on both. Scorers can then judge the final turn, every turn, or the whole trajectory. To serve single-turn and multi-turn records from one dataset, normalise first: turns = question if isinstance(question, list) else [question].
input_key, for example {"question": ["...", "..."]}, so the task receives the list.
Tie the run to the agent
Set the subject so the run appears on that agent’s profile in the Atlan app, and so you can list its history:subject_kind is agent, harness, or skill, and subject_id must be an artifact of that kind you can read. The experiment does not record which version of the agent ran. Put the version, commit, model, and prompt revision in config, or pin them with a context manifest.
Cost, tokens, and latency
Cost and tokens live on the traces, not in the summary. Query them per experiment with trace stats. A scalar query returns one number per measure; a table query grouped byscore_name returns the mean of each score across the run’s score spans.
0. To enforce a budget, compare these numbers in your CI script. The CI action reports them for information and never fails a check on them.
Nondeterminism
Agents vary run to run.Eval runs each case once. To measure variance, run the same eval several times with a trial field in config, and compare the means:
0 where your agent allows it. A gate threshold narrower than the run-to-run spread fails at random. Widen it, or gate on more cases.