> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Scores from your application

> Attach scores to production traces from your own code: inline checks while a request runs, and user feedback added after it finished.

Some signals only exist in production: a user's thumbs-up, a support agent's correction, a guardrail that fired. Record them as scores on the request's trace. They sit next to that request's latency, cost, and tool calls, and trace statistics aggregate them by score name.

This is separate from `Eval`, which writes scores for you, and from [hosted judges](/evals/online), which store their answers as insights.

## 1. Register the scorer once

Every score cites a scorer ID and version. Create a `code` scorer for an automated check, or a `human` scorer for feedback. Do this once, for example at deploy time, and keep the ID and `version_ordinal`. Scorer names are unique in the workspace, so the helper below reuses any scorer that already has the name; pick names specific to your application.

<CodeGroup>
  ```python Python theme={null}
  import os
  from atlanai import AtlanClient

  WORKSPACE = os.environ["ATLAN_WORKSPACE_ID"]
  client = AtlanClient(os.environ.get("ATLAN_BASE_URL", "https://api.atlan.com"),
                       bearer_token=os.environ["ATLAN_API_KEY"], workspace=WORKSPACE)


  def scorer(name: str, kind: str):
      for s in client.scorers.list(limit=500).items:
          if s.name == name:
              return s
      return client.scorers.create({"workspace_id": WORKSPACE, "name": name, "scorer_kind": kind})


  format_check = scorer("answer-format-check", "code")
  feedback = scorer("user-feedback", "human")
  ```

  ```typescript TypeScript theme={null}
  import { AtlanClient, rawEval } from "@atlanai/sdk";

  const workspace = process.env.ATLAN_WORKSPACE_ID!;
  const client = new AtlanClient({
    gatewayOrigin: process.env.ATLAN_BASE_URL ?? "https://api.atlan.com",
    bearerToken: process.env.ATLAN_API_KEY!,
    workspace,
  });

  async function scorer(name: string, kind: rawEval.EvalScorerKind) {
    const existing = (await client.scorers.list({ limit: 500 })).items.find((s) => s.name === name);
    return existing ?? client.scorers.create({ workspaceId: workspace, name, scorerKind: kind });
  }

  const formatCheck = await scorer("answer-format-check", rawEval.EvalScorerKind.Code);
  const feedback = await scorer("user-feedback", rawEval.EvalScorerKind.Human);
  ```
</CodeGroup>

A score without a valid scorer ID and version is **dropped with a warning, not raised**, so scoring never breaks a request. Set `ATLAN_DEBUG=true` to see the warnings while you wire this up.

## 2. Score while the request runs

`score` attaches a score to the span you call it on. Call it on the root span to score the whole request, or on a child span to score one step, such as a retrieval's precision. In Python, `score_trace` on any span scores the root. `init_logger`'s `project_name` is the service name on your traces; there is nothing to create first.

<CodeGroup>
  ```python Python theme={null}
  from atlanai.tracing import init_logger

  tracer = init_logger(project_name="support-agent").client

  with tracer.start_as_current_span("handle-request", as_type="task") as root:
      answer = your_agent("Where is my refund?")
      root.update(input="Where is my refund?", output=answer)
      root.score("cites_a_date", value=any(c.isdigit() for c in answer), data_type="BOOLEAN",
                 scorer_id=format_check.id, scorer_version=format_check.version_ordinal)
      trace_id, root_span_id = root.trace_id, root.span_id   # keep these to attach feedback later
  ```

  ```typescript TypeScript theme={null}
  import { initLogger } from "@atlanai/sdk/tracing";

  const logger = initLogger({ projectName: "support-agent" });
  const tracer = logger.client;
  let traceId = "", rootSpanId = "";   // keep these to attach feedback later

  await tracer.startAsCurrentSpan("handle-request", { asType: "task" }, async (root) => {
    const answer = await yourAgent("Where is my refund?");
    root.update({ input: "Where is my refund?", output: answer });
    root.score("cites_a_date", /\d/.test(answer), {
      dataType: "BOOLEAN", scorerId: formatCheck.id, scorerVersion: formatCheck.versionOrdinal,
    });
    traceId = root.traceId;
    rootSpanId = root.spanId;
  });
  ```
</CodeGroup>

| `data_type` | Pass | Stored as |
| - | - | - |
| `NUMERIC` (default) | A number | The number |
| `BOOLEAN` | `True`/`False` | `1.0`/`0.0` |
| `CATEGORICAL` | `string_value="approve"` | The label, with value `0.0` unless you pass one. Labels are kept on the span, but only numeric values are aggregated. |

Pass `comment="..."` with any score to record why.

## 3. Attach feedback after the request finished

Feedback often arrives minutes later, from another process. Keep the request's `trace_id` and root `span_id`, for example on your message record. Then open a span in that trace and score it:

<CodeGroup>
  ```python Python theme={null}
  with tracer.start_as_current_span("user-feedback", as_type="function",
          trace_context={"trace_id": trace_id, "parent_span_id": root_span_id}) as span:
      span.score("thumbs_up", value=1.0, comment="Clear and correct",
                 scorer_id=feedback.id, scorer_version=feedback.version_ordinal)
  tracer.flush()
  ```

  ```typescript TypeScript theme={null}
  await tracer.startAsCurrentSpan("user-feedback",
    { asType: "function", traceContext: { traceId, parentSpanId: rootSpanId } },
    async (span) => {
      span.score("thumbs_up", 1.0, {
        comment: "Clear and correct", scorerId: feedback.id, scorerVersion: feedback.versionOrdinal,
      });
    });
  await logger.flush();
  ```
</CodeGroup>

Use the IDs you kept from the original request. A `trace_context` with an invented trace ID creates a new, orphaned trace instead.

The feedback span becomes a child of the original root, so the request's trace shows the answer and the verdict together.

## Read scores back

Each score is a `score.<name>` span in the request's trace, a child of the span it was attached to, with `score_value`, `scorer_id`, and `scorer_version` set. Scores show on the trace in the Atlan app, and trace statistics aggregate `score_value` by `score_name`, alongside cost and latency. When the request belongs to a session, read the spans back with `client.sessions.traces.list_spans(session_id, trace_id, fields="core,attributes")`, or in TypeScript `client.sessions.traces.listSpans({ sessionId, traceId, fields: "core,attributes" })`. Spans become readable a few seconds after `flush()`, so poll briefly; see [make app traffic scoreable](/evals/online#make-app-traffic-scoreable). When a production score finds a failure worth keeping, [turn it into a test case](/evals/cookbooks/production-to-dataset).

## Keep scores consistent

* **Stable, lowercase names** with underscores, such as `thumbs_up` and `cites_a_date`. Aggregation is by name.
* **One meaning per scorer.** When a check's logic changes, register a new scorer name, or edit it through `PATCH /eval/v1/scorers/{id}` to record a new version and cite the new `version_ordinal`.
* **No user content in comments** beyond what your data policy allows. Comments are stored on the trace.
