Eval, which writes scores for you, and from hosted judges, which store their answers as insights.
1. Register the scorer once
Every score cites a scorer ID and version. Create acode scorer for an automated check, or a human scorer for feedback. Do this once, for example at deploy time, and keep the ID and version_ordinal. Scorer names are unique in the workspace, so the helper below reuses any scorer that already has the name; pick names specific to your application.
ATLAN_DEBUG=true to see the warnings while you wire this up.
2. Score while the request runs
score attaches a score to the span you call it on. Call it on the root span to score the whole request, or on a child span to score one step, such as a retrieval’s precision. In Python, score_trace on any span scores the root. init_logger’s project_name is the service name on your traces; there is nothing to create first.
Pass
comment="..." with any score to record why.
3. Attach feedback after the request finished
Feedback often arrives minutes later, from another process. Keep the request’strace_id and root span_id, for example on your message record. Then open a span in that trace and score it:
trace_context with an invented trace ID creates a new, orphaned trace instead.
The feedback span becomes a child of the original root, so the request’s trace shows the answer and the verdict together.
Read scores back
Each score is ascore.<name> span in the request’s trace, a child of the span it was attached to, with score_value, scorer_id, and scorer_version set. Scores show on the trace in the Atlan app, and trace statistics aggregate score_value by score_name, alongside cost and latency. When the request belongs to a session, read the spans back with client.sessions.traces.list_spans(session_id, trace_id, fields="core,attributes"), or in TypeScript client.sessions.traces.listSpans({ sessionId, traceId, fields: "core,attributes" }). Spans become readable a few seconds after flush(), so poll briefly; see make app traffic scoreable. When a production score finds a failure worth keeping, turn it into a test case.
Keep scores consistent
- Stable, lowercase names with underscores, such as
thumbs_upandcites_a_date. Aggregation is by name. - One meaning per scorer. When a check’s logic changes, register a new scorer name, or edit it through
PATCH /eval/v1/scorers/{id}to record a new version and cite the newversion_ordinal. - No user content in comments beyond what your data policy allows. Comments are stored on the trace.