> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Add scores

> Attach numeric, categorical, and boolean evaluation results to a trace or a single span.

A score records how well a run went. Attaching it through the SDK puts quality
next to latency and cost on the same trace, so you can query all three
together instead of joining an external evaluation store.

## Score a trace or a span

`score_trace` / `scoreTrace` attaches the score to the trace's root span. Use
it for a judgement about the whole run.

`score` attaches to the span you call it on. Use it when the judgement is about
one step — a retrieval's precision, a single tool call's correctness.

Every score must cite a Registry scorer artifact and immutable version. `Eval`
resolves or creates that scorer automatically. For manual tracing, resolve the
scorer once when the process starts and keep its `id` and `version_ordinal`.

<Warning>
  Automatic first-run registration cannot detect a change to a local scorer
  function or judge prompt. Give a materially changed scorer a new Registry
  name and pass its new ID and version. Do not patch a scorer definition and
  assume the same version ordinal now represents immutable evidence.
</Warning>

<CodeGroup>
  ```python Python theme={null}
  precision_scorer_id = "scorer_01example"
  precision_scorer_version = 3
  review_scorer_id = "scorer_02example"
  review_scorer_version = 5

  with client.start_as_current_span("review-pr", as_type="task") as root:
      with client.start_as_current_span("retrieve", as_type="tool") as retrieval:
          retrieval.score(
              "precision",
              value=0.8,
              scorer_id=precision_scorer_id,
              scorer_version=precision_scorer_version,
          )

      root.score_trace(
          "review_score",
          value=0.87,
          scorer_id=review_scorer_id,
          scorer_version=review_scorer_version,
      )
  ```

  ```typescript TypeScript theme={null}
  const precisionScorer = { id: "scorer_01example", version: 3 };
  const reviewScorer = { id: "scorer_02example", version: 5 };

  await client.startAsCurrentSpan("review-pr", { asType: "task" }, async (root) => {
    await client.startAsCurrentSpan("retrieve", { asType: "tool" }, async (retrieval) => {
      retrieval.score("precision", 0.8, {
        scorerId: precisionScorer.id,
        scorerVersion: precisionScorer.version,
      });
    });

    root.scoreTrace("review_score", 0.87, {
      scorerId: reviewScorer.id,
      scorerVersion: reviewScorer.version,
    });
  });
  ```
</CodeGroup>

`score_trace` targets the root span only when the SDK opened that root in the
current context. Otherwise it falls back to the span it was called on, so the
score is never silently dropped.

## The three data types

| Data type | Requires | Stored value |
| - | - | - |
| `NUMERIC` (default) | a numeric `value` | the number as given |
| `CATEGORICAL` | `string_value` / `stringValue` | the label, with numeric `0.0` unless you pass one |
| `BOOLEAN` | a truthy or falsy `value` | `1.0` or `0.0` |

<CodeGroup>
  ```python Python theme={null}
  identity = {"scorer_id": review_scorer_id, "scorer_version": review_scorer_version}

  root.score_trace("review_score", value=0.87, **identity)
  root.score_trace(
      "review_decision",
      data_type="CATEGORICAL",
      string_value="approve",
      **identity,
  )
  root.score_trace("blocking_present", value=False, data_type="BOOLEAN", **identity)
  ```

  ```typescript TypeScript theme={null}
  const identity = {
    scorerId: reviewScorer.id,
    scorerVersion: reviewScorer.version,
  };

  root.scoreTrace("review_score", 0.87, identity);
  root.scoreTrace("review_decision", undefined, {
    ...identity,
    dataType: "CATEGORICAL",
    stringValue: "approve",
  });
  root.scoreTrace("blocking_present", false, { ...identity, dataType: "BOOLEAN" });
  ```
</CodeGroup>

Pass `comment` on any score to record why the value was assigned.

## Invalid scores are dropped, not raised

Scoring never breaks the run. A score is dropped with a warning when the scorer
ID is absent or malformed, the scorer version is absent or less than 1, the
name is empty or not a string, the data type is unrecognized, a `NUMERIC` value
is not numeric, a `CATEGORICAL` score has no `string_value`, or a `BOOLEAN`
score has no value at all.

That means a typo in a data type produces a missing score rather than an
exception. If a score you expect is absent, enable `ATLAN_DEBUG=true` and check
the warnings.

## How a score is stored

Each score is a dedicated, immediately ended child span named `score.{name}`.
It carries the numeric value, data type, scorer artifact ID, scorer version,
and optional string value and comment as `atlan.score.*` attributes. This is
why one evaluated case can legitimately contain several score spans.

The score span's parent defines its scope. `score` parents it to the span you
call it on. `score_trace` parents it to the in-process root when that root can
be resolved. Use a stable, lowercase score name with underscores so reporting
aggregates the same measure across runs.

## Where scores surface

Scores are written for the reporting surfaces, which is where you read them
back alongside cost and latency — see [Reporting](/registry/reporting). The
management client's trace and span reads return execution shape and usage
totals rather than score values, so use reporting rather than
`registry_list_user_trace_spans` when you want to compare scores across runs.

## Next steps

<CardGroup cols={2}>
  <Card title="Trace your agent" icon="diagram-project" href="/tools/sdk/how-tos/tracing">
    Spans, observation types, usage and cost.
  </Card>

  <Card title="Reporting" icon="chart-line" href="/registry/reporting">
    Where scores surface alongside cost and latency.
  </Card>
</CardGroup>
