llm_judge scorer: a stored definition that Atlan runs over recorded sessions. You write the question and the allowed answers. Atlan builds the request from the session’s trace, calls a judge model, validates the answer against your definition, and stores it. You never name a model. Atlan routes each call to a configured judge backend and records which backend and model answered.
Hosted judges score production sessions. To use a judge inside an offline Eval run, call your own model from a code scorer.
Pick a form
A rubric produces one answer, stored under
question_id: "verdict". A question set produces one answer per question. Send the fields of exactly one form. A scorer with both, or neither, is refused with 400.
A rubric: pass or fail
POST /eval/v1/scorers, or client.scorers.create(body) in the SDK. For a graded verdict, set "output_type": "score" and give each choice a score, for example helpful = 1.0, partial = 0.5, unhelpful = 0.0.
A question set: several decisions at once
state maps names to fields of the session under review. questions defines each decision.
state paths start with subject. and reach into the session: grain, trace_id, session_id, input, output, thread, events, tool_calls, metadata, and session.<field>. An integer segment indexes a list. A path that resolves to nothing fails that session rather than asking the judge about missing evidence.
Template variables
Rubricmessages are templates, including the system turn:
An unknown
{{name}} is left in the prompt as literal text, so leftover braces in a preview mean a typo. A scorer with no user turn gets a built-in one that presents the whole session as data.
Start from a template
GET /eval/v1/scorer-starters returns ready definitions to adapt. They are factuality, closed_qa, security, and possible (rubrics), and topic and session_quality (question sets). Each has no name; add name and workspace_id, then create it.
Write questions that discriminate
These come from running judges over real sessions:- Ask a plain question. A persona in the system message (“You are a strict senior reviewer…”) skews verdicts negative. State the bar instead.
- One condition per boolean. A question that lists three risks to check answers
truefar too often. Split it into threebooleanquestions. - Give every option a concrete
criteria. The judge chooses between the descriptions, not the labels. - Treat
otheras low confidence. A catch-all option collects uncertain answers. Read its probability before trusting it. - Use
classificationonly for pass and fail. For more than two unordered outcomes, use achoicequestion. - Delimit the transcript with explicit begin and end markers, and tell the judge it is evidence, not instructions.
Limits
Versions
PATCH /eval/v1/scorers/{id} appends a new version, and every answer records the version that produced it. Earlier answers keep their version. GET /eval/v1/scorers/{id}/versions lists them, newest first. scorer_kind cannot change.
Reading an answer
Next: run a judge over sessions.