Skip to main content
Online evals score sessions your agent already produced. A stored llm_judge scorer defines the question. Agent Gateway reads the selected session’s trace, sends a bounded transcript to a hosted judge, and writes one session_insight per answer. The judge is one model call; it does not run your agent or call its tools.
SDK version prerequisite. The SDK examples below require atlanai or @atlanai/sdk 0.3.5 or later. The curl examples call the Gateway route directly. The public contract can start a scoring run, but its status and insight reads, and the action route for automatic scoring, are private. A 202 response means the run was queued, not that scoring finished.

Before you start

You need a completed or failed session with a readable trace, a workspace-scoped bearer token, and builder or admin access to the scorer’s workspace. A scorer in another workspace cannot judge that session. The deployment must also have a judge backend configured; otherwise the score route returns 503. The judge receives session content. Apply tracing privacy controls before you collect real traffic. mode: "preview" returns the exact judge request to the caller, so inspect or log it only in an approved environment. Set these values in your shell using an approved secret store:

Create a rubric scorer

A rubric produces one verdict, stored under question_id: "verdict". The example asks whether the final answer addresses the user’s request. classification requires the choices pass and fail.
Copy the returned scorer_... ID to ATLAN_SCORER_ID. There is no model field on a scorer: the gateway chooses a configured backend for each call and records backend and model on the answer. Editing a scorer appends a version; existing answers keep the version that produced them. For several independent answers, use the question-set form instead: state maps names to subject. paths and questions defines each decision. Do not combine question-set fields with messages, output_type, or choices in one scorer.
Each question becomes its own insight row. A state path that resolves to no field fails the subject rather than asking the judge to score missing evidence. GET /eval/v1/scorer-starters offers example definitions for both forms.

Preview, then test

Preview builds the judge request but makes no model call and writes no insight. Test calls the judge and returns its answers without writing insights. Both are synchronous and accept at most five subjects.
Change mode to test after checking the previewed evidence. Test mode incurs a judge call. A session with no trace or an empty readable transcript receives an error line and no verdict.

Start a scoring run

score is the default mode. The route resolves the sessions first, then returns 202 with a score_run_id. Judging proceeds on the action worker.
The accepted response includes score_run_id, scorer_version, run_status: "queued", total, grain, and concurrency. It does not contain verdicts. The run row later records queued, running, and then succeeded or failed, along with done, failed, and a per-subject report. A completed run may have subject failures; inspect each line rather than treating succeeded as a perfect score. The current public contract has no typed read route for score_run or session_insight. Until one is published, the public SDK can start a run but cannot retrieve its progress or durable answers. The same visibility gap prevents a complete public recipe for automatic session.created scoring. Operators with approved internal access can inspect those artifacts; public clients should wait for the supported read surface.

Use the generated SDK

Install SDK 0.3.5 or later for score_with in Python or scoreWith in TypeScript. The SDK sends the workspace header when workspace is set, but workspace_id or workspaceId is still required in a scorer create body.

Choose the subjects

Send exactly one selector family per call. The named family may combine session_ids and trace_ids; the others cannot be combined with it or each other. The route accepts at most 100 subjects in score mode and five in preview or test. The default grain: "trace" judges one run; grain: "session" reads every trace of the session as one transcript, which can cost more. limit defaults to 20 for list selectors. Keep concurrency at its default of 4 until a run’s measured elapsed_ms and rate_limited results justify a change. The offline guide runs against the current SDK without the hosted scoring route.