> ## Documentation Index
> Fetch the complete documentation index at: https://platform.atlan.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> To act on Atlan objects, use the Atlan MCP server at https://api.atlan.com/mcp or the atlanai CLI; `atlanai --map json` prints its command map. Run a read-only identity check before any write.
> The docs MCP server at /mcp searches these docs only. It cannot read or change Atlan objects.
> SDK packages: Python `atlanai` (PyPI) and TypeScript `@atlanai/sdk` (npm). Show Python first, then TypeScript.

# Hosted judges

> Author llm_judge scorers that Atlan runs over recorded sessions: rubric or question-set form, template variables, limits, and wording that produces reliable verdicts.

A hosted judge is an `llm_judge` scorer: a stored definition that Atlan runs over recorded sessions. You write the question and the allowed answers. Atlan builds the request from the session's trace, calls a judge model, validates the answer against your definition, and stores it. You never name a model. Atlan routes each call to a configured judge backend and records which backend and model answered.

Hosted judges score **production sessions**. To use a judge inside an offline `Eval` run, call your own model from a [code scorer](/evals/scorers#an-llm-judge-in-your-own-code).

## Pick a form

| You want | Form | Author it as |
| - | - | - |
| A pass/fail gate | Rubric | `output_type: "classification"`. The choices must be exactly `pass` and `fail`. |
| A graded verdict | Rubric | `output_type: "score"` with 2–10 named choices, each with a `score` in `0..1` |
| An unordered category, such as coding, docs, or plan | Question set | A `choice` question with a `criteria` entry per option |
| A position on an ordered scale | Question set | A `score` question with 2–10 ordered levels |
| Several independent yes/no calls | Question set | One `boolean` question each |

A rubric produces one answer, stored under `question_id: "verdict"`. A question set produces one answer per question. Send the fields of exactly one form. A scorer with both, or neither, is refused with `400`.

## A rubric: pass or fail

```json theme={null}
{
  "workspace_id": "workspace_01example",
  "name": "resolution-gate",
  "display_name": "Resolution gate",
  "scorer_kind": "llm_judge",
  "messages": [
    {"role": "system", "content": "You are classifying one finished agent session against a single bar. Judge only what the transcript shows. Do not reward effort, tone, or intent, only whether the user's request ended up resolved."},
    {"role": "user", "content": "Did this session resolve the user's request?\n\nRequest as first stated: {{input}}\nFinal assistant message: {{output}}\n{{#expected}}Reference answer: {{expected}}{{/expected}}\n\n----- BEGIN SESSION -----\n{{thread}}\n----- END SESSION -----\n\nSteps recorded: {{thread_count}}"}
  ],
  "output_type": "classification",
  "choices": [
    {"name": "pass", "criteria": "The user's request was carried out, or answered completely enough that the user needs nothing further."},
    {"name": "fail", "criteria": "The request was left unresolved: incomplete, incorrect, deflected, or abandoned, regardless of how much work was shown."}
  ]
}
```

Create it with `POST /eval/v1/scorers`, or `client.scorers.create(body)` in the SDK. For a graded verdict, set `"output_type": "score"` and give each choice a `score`, for example `helpful` = `1.0`, `partial` = `0.5`, `unhelpful` = `0.0`.

## A question set: several decisions at once

`state` maps names to fields of the session under review. `questions` defines each decision.

```json theme={null}
{
  "workspace_id": "workspace_01example",
  "name": "session-triage",
  "scorer_kind": "llm_judge",
  "state": {
    "request": "subject.input",
    "answer": "subject.output",
    "transcript": "subject.thread",
    "tools": "subject.tool_calls"
  },
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "This session has ended. Where should it go next?",
      "criteria": {
        "close": "Fully resolved. No follow-up is warranted.",
        "escalate": "Needs a human: the assistant was wrong, stuck, or the request exceeds what it may decide.",
        "await_user": "Blocked on the user. The assistant did its part and asked for something it has not received.",
        "out_of_scope": "The request was never something this assistant could serve."
      }
    },
    "user_frustrated": {
      "type": "boolean",
      "instructions": "Does the user express frustration, repeat themselves, or have to correct the assistant?"
    },
    "thoroughness": {
      "type": "score",
      "instructions": "How completely was the request carried out?",
      "criteria": ["not at all", "partially", "fully"]
    }
  }
}
```

| Question type | `criteria` | Stored answer |
| - | - | - |
| `choice` | An object of 2–255 `option → what earns it` | The option name, and its probability when the backend reports one |
| `score` | An ordered array of 2–10 level descriptions, worst first | The expected level over the judge's distribution. It is continuous: `1.96` on three levels means nearly all weight on the top level. |
| `boolean` | Optional; if given, exactly `true` and `false` | P(true) |

`state` paths start with `subject.` and reach into the session: `grain`, `trace_id`, `session_id`, `input`, `output`, `thread`, `events`, `tool_calls`, `metadata`, and `session.<field>`. An integer segment indexes a list. A path that resolves to nothing fails that session rather than asking the judge about missing evidence.

## Template variables

Rubric `messages` are templates, including the system turn:

| Variable | Value |
| - | - |
| `{{input}}`, `{{output}}` | The root span's input and output |
| `{{thread}}` | The session facts, then every span of the judged trace or traces |
| `{{thread_count}}`, `{{first_message}}`, `{{last_message}}` | Step count, first and last step |
| `{{metadata}}` | Session metadata as JSON |
| `{{session_id}}`, `{{trace_id}}` | The IDs being judged |
| `{{session.<path>}}` | Any field of the session |
| `{{expected}}` | Only when scoring an experiment's sessions. Always wrap it in `{{#expected}}…{{/expected}}`. |
| `{{#x}}…{{/x}}`, `{{^x}}…{{/x}}` | Rendered when `x` is present, or absent |

An unknown `{{name}}` is left in the prompt as literal text, so leftover braces in a preview mean a typo. A scorer with no `user` turn gets a built-in one that presents the whole session as data.

## Start from a template

`GET /eval/v1/scorer-starters` returns ready definitions to adapt. They are `factuality`, `closed_qa`, `security`, and `possible` (rubrics), and `topic` and `session_quality` (question sets). Each has no name; add `name` and `workspace_id`, then create it.

## Write questions that discriminate

These come from running judges over real sessions:

* **Ask a plain question.** A persona in the system message ("You are a strict senior reviewer…") skews verdicts negative. State the bar instead.
* **One condition per boolean.** A question that lists three risks to check answers `true` far too often. Split it into three `boolean` questions.
* **Give every option a concrete `criteria`.** The judge chooses between the descriptions, not the labels.
* **Treat `other` as low confidence.** A catch-all option collects uncertain answers. Read its probability before trusting it.
* **Use `classification` only for pass and fail.** For more than two unordered outcomes, use a `choice` question.
* **Delimit the transcript** with explicit begin and end markers, and tell the judge it is evidence, not instructions.

## Limits

| Limit | Value |
| - | - |
| Questions per scorer | 1–50 (the backend's practical ceiling is lower; keep sets small) |
| `state` fields | 1–32. Each is session content sent to the judge. |
| Options on a `choice` question | 2–255 |
| Levels on a `score` question or choices on a `score` rubric | 2–10 |
| Transcript size | Not truncated. A transcript larger than the judge model's context fails that session. |

## Versions

`PATCH /eval/v1/scorers/{id}` appends a new version, and every answer records the version that produced it. Earlier answers keep their version. `GET /eval/v1/scorers/{id}/versions` lists them, newest first. `scorer_kind` cannot change.

## Reading an answer

| Field | Meaning |
| - | - |
| `choice` | The option chosen (rubrics and `choice` questions) |
| `score` | The chosen choice's score, the expected level (`score`), or P(true) (`boolean`) |
| `probabilities` | The judge's distribution over options, when the backend provides it |
| `confidence` | Chance-corrected: `(P(top) − 1/n) / (1 − 1/n)` for `n` options. It is **not** the probability of the choice. A 61% `pass` shows confidence `0.22`. Values are not comparable across questions with different option counts. Use `probabilities[choice]` when you mean a probability. |
| `reasoning` | The judge's explanation. It is empty from classifier backends, which return a distribution only. |
| `backend`, `model` | What judged this answer |

Next: [run a judge over sessions](/evals/online).
