Skip to main content
Some signals only exist in production: a user’s thumbs-up, a support agent’s correction, a guardrail that fired. Record them as scores on the request’s trace. They sit next to that request’s latency, cost, and tool calls, and trace statistics aggregate them by score name. This is separate from Eval, which writes scores for you, and from hosted judges, which store their answers as insights.

1. Register the scorer once

Every score cites a scorer ID and version. Create a code scorer for an automated check, or a human scorer for feedback. Do this once, for example at deploy time, and keep the ID and version_ordinal. Scorer names are unique in the workspace, so the helper below reuses any scorer that already has the name; pick names specific to your application.
A score without a valid scorer ID and version is dropped with a warning, not raised, so scoring never breaks a request. Set ATLAN_DEBUG=true to see the warnings while you wire this up.

2. Score while the request runs

score attaches a score to the span you call it on. Call it on the root span to score the whole request, or on a child span to score one step, such as a retrieval’s precision. In Python, score_trace on any span scores the root. init_logger’s project_name is the service name on your traces; there is nothing to create first.
Pass comment="..." with any score to record why.

3. Attach feedback after the request finished

Feedback often arrives minutes later, from another process. Keep the request’s trace_id and root span_id, for example on your message record. Then open a span in that trace and score it:
Use the IDs you kept from the original request. A trace_context with an invented trace ID creates a new, orphaned trace instead. The feedback span becomes a child of the original root, so the request’s trace shows the answer and the verdict together.

Read scores back

Each score is a score.<name> span in the request’s trace, a child of the span it was attached to, with score_value, scorer_id, and scorer_version set. Scores show on the trace in the Atlan app, and trace statistics aggregate score_value by score_name, alongside cost and latency. When the request belongs to a session, read the spans back with client.sessions.traces.list_spans(session_id, trace_id, fields="core,attributes"), or in TypeScript client.sessions.traces.listSpans({ sessionId, traceId, fields: "core,attributes" }). Spans become readable a few seconds after flush(), so poll briefly; see make app traffic scoreable. When a production score finds a failure worth keeping, turn it into a test case.

Keep scores consistent

  • Stable, lowercase names with underscores, such as thumbs_up and cites_a_date. Aggregation is by name.
  • One meaning per scorer. When a check’s logic changes, register a new scorer name, or edit it through PATCH /eval/v1/scorers/{id} to record a new version and cite the new version_ordinal.
  • No user content in comments beyond what your data policy allows. Comments are stored on the trace.