Skip to main content
A hosted judge is an llm_judge scorer: a stored definition that Atlan runs over recorded sessions. You write the question and the allowed answers. Atlan builds the request from the session’s trace, calls a judge model, validates the answer against your definition, and stores it. You never name a model. Atlan routes each call to a configured judge backend and records which backend and model answered. Hosted judges score production sessions. To use a judge inside an offline Eval run, call your own model from a code scorer.

Pick a form

A rubric produces one answer, stored under question_id: "verdict". A question set produces one answer per question. Send the fields of exactly one form. A scorer with both, or neither, is refused with 400.

A rubric: pass or fail

Create it with POST /eval/v1/scorers, or client.scorers.create(body) in the SDK. For a graded verdict, set "output_type": "score" and give each choice a score, for example helpful = 1.0, partial = 0.5, unhelpful = 0.0.

A question set: several decisions at once

state maps names to fields of the session under review. questions defines each decision.
state paths start with subject. and reach into the session: grain, trace_id, session_id, input, output, thread, events, tool_calls, metadata, and session.<field>. An integer segment indexes a list. A path that resolves to nothing fails that session rather than asking the judge about missing evidence.

Template variables

Rubric messages are templates, including the system turn: An unknown {{name}} is left in the prompt as literal text, so leftover braces in a preview mean a typo. A scorer with no user turn gets a built-in one that presents the whole session as data.

Start from a template

GET /eval/v1/scorer-starters returns ready definitions to adapt. They are factuality, closed_qa, security, and possible (rubrics), and topic and session_quality (question sets). Each has no name; add name and workspace_id, then create it.

Write questions that discriminate

These come from running judges over real sessions:
  • Ask a plain question. A persona in the system message (“You are a strict senior reviewer…”) skews verdicts negative. State the bar instead.
  • One condition per boolean. A question that lists three risks to check answers true far too often. Split it into three boolean questions.
  • Give every option a concrete criteria. The judge chooses between the descriptions, not the labels.
  • Treat other as low confidence. A catch-all option collects uncertain answers. Read its probability before trusting it.
  • Use classification only for pass and fail. For more than two unordered outcomes, use a choice question.
  • Delimit the transcript with explicit begin and end markers, and tell the judge it is evidence, not instructions.

Limits

Versions

PATCH /eval/v1/scorers/{id} appends a new version, and every answer records the version that produced it. Earlier answers keep their version. GET /eval/v1/scorers/{id}/versions lists them, newest first. scorer_kind cannot change.

Reading an answer

Next: run a judge over sessions.