> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Evaluate an agent

> Wrap an agent as an eval task, trace its model and tool calls, and score tool choice, arguments, trajectory, multi-turn behaviour, and cost.

An agent can reach the right answer by the wrong path, or the wrong answer by a reasonable one. Score both: the **final output**, and the **trajectory**, meaning which tools it called, with which arguments, in how many steps.

The pattern has three parts:

1. **The task returns the trajectory, not just the answer.** Your scorers see only what the task returns, so return the tool calls alongside the answer.
2. **The trace records everything else.** Model calls, tool spans, tokens, cost, and latency land on the case's trace for diagnosis, without you returning them.
3. **The dataset states what a good trajectory is.** Put the expected tools and arguments in `expected`.

## Wrap the agent as a task

Give your agent a thin adapter that returns a structured output: the answer, plus the tool calls the agent made.

Record the tool calls **at the tool boundary**, where they execute, not from what the agent says it did. A model's own account of its calls can differ from what ran. Here each tool is traced and records its own calls into a per-case list. The list is a context variable, so concurrent cases do not mix their calls.

<CodeGroup>
  ```python Python theme={null}
  import contextvars
  import functools
  import inspect
  import re

  from atlanai.tracing import auto_instrument, init_logger, observe

  logger = init_logger(project_name="order-agent")   # the service name on traces; nothing to create first
  auto_instrument()  # enable installed model and framework instrumentors

  _tool_calls = contextvars.ContextVar("tool_calls", default=None)


  def recorded(fn):
      """Record each call the agent actually makes: tool name and bound arguments."""
      signature = inspect.signature(fn)

      @functools.wraps(fn)
      def wrapper(*args, **kwargs):
          calls = _tool_calls.get()
          if calls is not None:
              calls.append({"name": fn.__name__, "args": dict(signature.bind(*args, **kwargs).arguments)})
          return fn(*args, **kwargs)
      return wrapper


  @observe(as_type="tool")
  @recorded
  def lookup_order(order_id: str) -> dict:
      return {"order_id": order_id, "status": "shipped", "eta": "Friday"}


  def order_agent(message: str, history: list[dict]) -> str:
      # Stand-in for your agent: it decides which tools to call.
      text = " ".join([m["content"] for m in history if m["role"] == "user"] + [message])
      order_id = re.search(r"\d{4}", text).group()
      order = lookup_order(order_id)
      return f"Order {order_id} arrives {order['eta']}."


  def order_task(question: str) -> dict:
      calls = []
      token = _tool_calls.set(calls)
      try:
          answer = order_agent(question, [])
      finally:
          _tool_calls.reset(token)
      return {"answer": answer, "tool_calls": calls}
  ```

  ```typescript TypeScript theme={null}
  import { AsyncLocalStorage } from "node:async_hooks";
  import { initLogger, wrapTraced } from "@atlanai/sdk/tracing";

  const logger = initLogger({ projectName: "order-agent" });   // the service name on traces

  type ToolCall = { name: string; args: Record<string, unknown> };
  const toolCalls = new AsyncLocalStorage<ToolCall[]>();

  // Trace the tool, and record each call the agent actually makes.
  function recorded<A extends Record<string, unknown>, R>(name: string, fn: (args: A) => Promise<R>) {
    return wrapTraced(async (args: A) => {
      toolCalls.getStore()?.push({ name, args });
      return fn(args);
    }, { name, type: "tool" });
  }

  const lookupOrder = recorded("lookup_order", async ({ order_id }: { order_id: string }) =>
    ({ order_id, status: "shipped", eta: "Friday" }));

  type Message = { role: "user" | "assistant"; content: string };

  async function orderAgent(message: string, history: Message[]): Promise<string> {
    // Stand-in for your agent: it decides which tools to call.
    const text = [...history.filter((m) => m.role === "user").map((m) => m.content), message].join(" ");
    const orderId = text.match(/\d{4}/)?.[0] ?? "";
    const order = await lookupOrder({ order_id: orderId });
    return `Order ${orderId} arrives ${order.eta}.`;
  }

  async function orderTask(question: string) {
    const calls: ToolCall[] = [];
    const answer = await toolCalls.run(calls, () => orderAgent(question, []));
    return { answer, toolCalls: calls };
  }
  ```
</CodeGroup>

`project_name` is the service name on your traces; there is nothing to create in Atlan first. Create the logger yourself when you instrument your agent, and pass it to `Eval` as `trace_logger` in Python or `logger` in TypeScript (see [run it](#run-it)). Its spans then nest under each case. Every case trace looks like this:

```text theme={null}
eval.case.1            task
├── eval.task          task       ← your agent runs here
│   ├── chat …         llm        ← from an instrumented model client
│   └── lookup_order   tool
├── scorer.tool_choice function
└── score.tool_choice  score      ← the score, citing the scorer version
```

Model and framework spans appear when their instrumentation is installed. In Python, `auto_instrument()` enables installed OpenTelemetry instrumentors for providers such as OpenAI and Anthropic, and frameworks such as LangChain, CrewAI, and the OpenAI Agents SDK. In TypeScript, wrap the Vercel AI SDK with `wrapAISDK`. See [tracing integrations](/tools/sdk/how-tos/integrations) for the full list. Every eval case is traced in full: `Eval` refuses a logger that samples or drops content.

## Stub tools with side effects

An eval runs your agent for real. A tool that sends email, issues a refund, or writes to production does that on every case of every run. Give the agent a stand-in for each such tool: it records the call, so the scorers still see it, and returns a plausible result without acting.

<CodeGroup>
  ```python Python theme={null}
  @observe(as_type="tool", name="issue_refund")
  @recorded
  def issue_refund(order_id: str, amount: float) -> dict:
      return {"refund_id": "refund_test", "status": "queued"}   # nothing is refunded
  ```

  ```typescript TypeScript theme={null}
  const issueRefund = recorded("issue_refund", async ({ order_id, amount }: { order_id: string; amount: number }) =>
    ({ refund_id: "refund_test", status: "queued", order_id, amount }));   // nothing is refunded
  ```
</CodeGroup>

Pass the stand-ins where your agent looks up its tools, for example its tool registry or constructor. Keep each stand-in's name and signature identical to the real tool's, so the tool schema the model sees does not change.

## Score tool calls

These scorers read `output["tool_calls"]` and `expected["tools"]` as lists of `{name, args}`; in TypeScript, `output.toolCalls`. Adapt the field names to your adapter.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import ScoreValue


  def tool_choice(*, output, expected, **_):
      """Did the agent call exactly the expected tools, in order?"""
      got = [c["name"] for c in output["tool_calls"]]
      want = [c["name"] for c in expected["tools"]]
      return ScoreValue(score=float(got == want), metadata={"called": got, "expected": want})


  def tool_recall(*, output, expected, **_):
      """Share of expected tools that were called, ignoring order and extras."""
      got = {c["name"] for c in output["tool_calls"]}
      want = {c["name"] for c in expected["tools"]}
      return len(got & want) / len(want) if want else None


  def tool_arguments(*, output, expected, **_):
      """Share of expected calls made with exactly those arguments. Each actual call matches once."""
      unmatched = list(output["tool_calls"])
      hits = 0
      for want in expected["tools"]:
          if want in unmatched:
              unmatched.remove(want)
              hits += 1
      return hits / len(expected["tools"]) if expected["tools"] else None


  def step_budget(*, output, expected, **_):
      """1.0 at or under the expected number of calls, falling off after."""
      budget = max(len(expected["tools"]), 1)
      return min(1.0, budget / max(len(output["tool_calls"]), 1))


  def answer_contains(*, output, expected, **_):
      return float(expected["answer_contains"].lower() in output["answer"].lower())
  ```

  ```typescript TypeScript theme={null}
  type Output = { answer: string; toolCalls: ToolCall[] };
  type Expected = { tools: ToolCall[]; answerContains: string };
  type Args = { output: Output; expected?: Expected };
  // Compare as JSON with object keys sorted: stored `expected` values do not keep key order.
  const canonical = (v: unknown): unknown =>
    Array.isArray(v) ? v.map(canonical)
    : v !== null && typeof v === "object"
      ? Object.fromEntries(Object.entries(v).sort(([a], [b]) => (a < b ? -1 : a > b ? 1 : 0)).map(([k, x]) => [k, canonical(x)]))
      : v;
  const same = (a: unknown, b: unknown) => JSON.stringify(canonical(a)) === JSON.stringify(canonical(b));

  const toolChoice = {
    name: "tool_choice",
    scorer: ({ output, expected }: Args) => {
      if (!expected) return null;
      const got = output.toolCalls.map((c) => c.name);
      const want = expected.tools.map((c) => c.name);
      return { score: Number(same(got, want)), metadata: { called: got, expected: want } };
    },
  };

  const toolRecall = {
    name: "tool_recall",
    scorer: ({ output, expected }: Args) => {
      if (!expected?.tools.length) return null;
      const got = new Set(output.toolCalls.map((c) => c.name));
      return expected.tools.filter((c) => got.has(c.name)).length / expected.tools.length;
    },
  };

  const toolArguments = {
    name: "tool_arguments",
    scorer: ({ output, expected }: Args) => {
      if (!expected?.tools.length) return null;
      const unmatched = [...output.toolCalls];
      let hits = 0;
      for (const want of expected.tools) {
        const i = unmatched.findIndex((c) => same(c, want));
        if (i >= 0) { unmatched.splice(i, 1); hits++; }
      }
      return hits / expected.tools.length;
    },
  };

  const stepBudget = {
    name: "step_budget",
    scorer: ({ output, expected }: Args) =>
      expected ? Math.min(1, Math.max(expected.tools.length, 1) / Math.max(output.toolCalls.length, 1)) : null,
  };

  const answerContains = {
    name: "answer_contains",
    scorer: ({ output, expected }: Args) =>
      expected ? Number(output.answer.toLowerCase().includes(expected.answerContains.toLowerCase())) : null,
  };
  ```
</CodeGroup>

In TypeScript, declare `expected` optional and return `null` when it is missing: `Eval` types it as possibly absent, and a required `expected` fails `tsc --strict`. In TypeScript, compare structured values with `same`, not raw `JSON.stringify`: `expected` read from a dataset comes back with its object keys in a different order, and a raw comparison then scores `0` on every case. `tool_arguments` matches each expected call to at most one actual call, so an agent that calls a tool once cannot satisfy two expected calls to it.

Which to use:

| Question | Scorer | When it matters |
| - | - | - |
| Did it pick the right tools? | `tool_choice` (strict order) or `tool_recall` (set) | Always. Use strict order only when order is part of correctness. |
| Did it call them correctly? | `tool_arguments` | When the wrong ID or filter silently returns a plausible wrong answer. |
| Did it take a sane path? | `step_budget` | When loops, retries, or redundant calls cost money or latency. |
| Is the final answer right? | `answer_contains`, [exact match, required facts](/evals/scorers#catalog), or an [LLM judge](/evals/scorers#an-llm-judge-in-your-own-code) | Always. The trajectory scores explain it. |
| Did it refuse or leak? | [Pattern checks](/evals/scorers#pattern-and-refusal-checks) | For user-facing agents |

Compare arguments on the fields that matter. Normalise or drop volatile ones, such as timestamps and request IDs, before comparing, or the score measures noise.

## Run it

Put the agent, the scorers, and this call in one file and run it.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import Eval

  run = Eval(
      "order-agent",
      data=[{
          "id": "eta",
          "input": "When will order 2210 arrive?",
          "expected": {"tools": [{"name": "lookup_order", "args": {"order_id": "2210"}}],
                       "answer_contains": "Friday"},
      }],
      task=order_task,
      scores=[tool_choice, tool_recall, tool_arguments, step_budget, answer_contains],
      trace_logger=logger,
  )
  print(run.summary["scores"])
  ```

  ```typescript TypeScript theme={null}
  import { Eval } from "@atlanai/sdk";

  const run = await Eval("order-agent", {
    data: () => [{
      id: "eta",
      input: "When will order 2210 arrive?",
      expected: { tools: [{ name: "lookup_order", args: { order_id: "2210" } }], answerContains: "Friday" },
    }],
    task: orderTask,
    scores: [toolChoice, toolRecall, toolArguments, stepBudget, answerContains],
  }, { logger });
  console.log(run.summary.scores);
  ```
</CodeGroup>

## Evaluate a multi-turn conversation

Put the user turns in `input` as a list. The task replays them against your agent, passing the history so far, and returns the transcript and every tool call across the turns. It returns the same `answer` and `tool_calls` shape as the single-turn task, with the final reply as `answer`, so the scorers above work on both. Scorers can then judge the final turn, every turn, or the whole trajectory. To serve single-turn and multi-turn records from one dataset, normalise first: `turns = question if isinstance(question, list) else [question]`.

<CodeGroup>
  ```python Python theme={null}
  from atlanai import Eval

  # Uses _tool_calls, logger, and order_agent from the adapter above.


  def conversation_task(turns: list[str]) -> dict:
      history, replies, calls = [], [], []
      token = _tool_calls.set(calls)
      try:
          for user_turn in turns:
              reply = order_agent(user_turn, history)   # your agent, with the history so far
              history += [{"role": "user", "content": user_turn}, {"role": "assistant", "content": reply}]
              replies.append(reply)
      finally:
          _tool_calls.reset(token)
      return {"answer": replies[-1], "replies": replies, "tool_calls": calls}


  def remembers_order_id(*, output, **_):
      # The order ID from turn 1 should still be used in the last reply.
      return float("2210" in output["answer"])


  run = Eval(
      "order-agent-multi-turn",
      data=[{"id": "follow-up", "input": ["Where is order 2210?", "And when will it arrive?"]}],
      task=conversation_task,
      scores=[remembers_order_id],
      trace_logger=logger,
  )
  print(run.summary["scores"])
  ```

  ```typescript TypeScript theme={null}
  async function conversationTask(turns: string[]) {
    const history: Message[] = [];
    const replies: string[] = [];
    const calls: ToolCall[] = [];
    await toolCalls.run(calls, async () => {
      for (const userTurn of turns) {
        const reply = await orderAgent(userTurn, history);   // your agent, with the history so far
        history.push({ role: "user", content: userTurn }, { role: "assistant", content: reply });
        replies.push(reply);
      }
    });
    return { answer: replies[replies.length - 1], replies, toolCalls: calls };
  }

  const multi = await Eval("order-agent-multi-turn", {
    data: () => [{ id: "follow-up", input: ["Where is order 2210?", "And when will it arrive?"] }],
    task: conversationTask,
    scores: [{ name: "remembers_order_id", scorer: ({ output }: { output: { answer: string } }) => Number(output.answer.includes("2210")) }],
  }, { logger });
  console.log(multi.summary.scores);
  ```
</CodeGroup>

In a dataset, store the turns under the `input_key`, for example `{"question": ["...", "..."]}`, so the task receives the list.

## Tie the run to the agent

Set the subject so the run appears on that agent's profile in the Atlan app, and so you can list its history:

<CodeGroup>
  ```python Python theme={null}
  run = Eval("order-agent", dataset="order-agent-cases", task=order_task,
             scores=[tool_choice, answer_contains], trace_logger=logger,
             subject_kind="agent", subject_id="agent_01example")

  history = client.experiments.list(subject_kind="agent", subject_id="agent_01example")
  ```

  ```typescript TypeScript theme={null}
  const run = await Eval("order-agent", { dataset: "order-agent-cases", task: orderTask, scores: [toolChoice, answerContains] },
    { logger, subjectKind: "agent", subjectId: "agent_01example" });
  ```
</CodeGroup>

`subject_kind` is `agent`, `harness`, or `skill`, and `subject_id` must be an artifact of that kind you can read. The experiment does not record *which version* of the agent ran. Put the version, commit, model, and prompt revision in `config`, or pin them with a [context manifest](/evals/offline#pin-the-context-that-changed).

## Cost, tokens, and latency

Cost and tokens live on the traces, not in the summary. Query them per experiment with trace stats. A scalar query returns one number per measure; a table query grouped by `score_name` returns the mean of each score across the run's score spans.

```python theme={null}
import time

now = int(time.time())
window = {"start_time_unix_seconds": now - 86_400, "end_time_unix_seconds": now}

totals = client.experiments.traces.query_stats(run.experiment_id, {
    **window, "request_type": "scalar",
    "queries": [{"name": "cost", "measure": "cost"},
                {"name": "tokens", "measure": "tokens"},
                {"name": "p90", "measure": "latency_p90"}],
})
for r in totals.data.results:
    print(r.query_name, r.value, r.unit)
```

Token and cost values come from model spans that report usage. An agent without instrumented model calls shows `0`. To enforce a budget, compare these numbers in your CI script. The [CI action](/evals/ci) reports them for information and never fails a check on them.

## Nondeterminism

Agents vary run to run. `Eval` runs each case once. To measure variance, run the same eval several times with a `trial` field in `config`, and compare the means:

```python theme={null}
means = []
for trial in range(3):
    r = Eval("order-agent", dataset="order-agent-cases", task=order_task,
             scores=[tool_choice], config={"model": "example-model", "trial": trial})
    means.append(r.summary["scores"]["tool_choice"]["mean"])
print(min(means), max(means))
```

Set model temperature to `0` where your agent allows it. A gate threshold narrower than the run-to-run spread fails at random. Widen it, or gate on more cases.
