Skip to main content
A code scorer is a function your eval runs on each case’s output. It returns a number between 0 and 1, where higher is better. Eval records the score on the case’s result and trace, and registers the scorer as a versioned artifact so every score says which definition produced it. Use code scorers for anything you can check deterministically: equality, structure, tool calls, keywords, budgets. Reach for an LLM judge only for what code cannot decide. For judges over production sessions, see hosted judges.

The contract

Arguments. Python matches parameters by keyword from input, output, expected, and metadata, so declare only what you use and accept **_ for the rest. TypeScript passes one object with input, output, expected, metadata, tags, and id. metadata is the inline case’s metadata, or the dataset record’s extra. In TypeScript, expected is optional on every case, so type it expected?: and return null when it is missing; a required expected fails tsc --strict. Passing scorers to Eval. Python accepts bare functions, named after the function, or {"name": ..., "scorer": fn} mappings. A lambda should always be wrapped in a mapping with a name. TypeScript accepts { name, scorer } objects only. Return values. A non-finite number raises. Two scorers that emit the same score name fail the run: every score name belongs to exactly one scorer. Failures. If a scorer raises, its message is appended to that case’s error and the other scorers still run. If the task raises, the case records the error and its scorers are skipped.

Versioning and names

Eval registers each scorer as <eval-name>-<scorer-name> in the workspace, lowercased with _ turned into -: scorer exact_match in eval support-agent is support-agent-exact-match. Score names on results keep the function’s own name, exact_match. Eval stores a digest of the function’s source. When the source changes, the next run records a new scorer version, and each score cites the version that produced it.
  • Keep the Eval name stable. Renaming the eval creates new scorer artifacts, so a baseline and a candidate with different eval names are not scored by the same scorer. The CI action also matches baselines on this name.
  • Share one scorer across evals by passing registry_name (registryName in TypeScript) in the scorer mapping.
  • Pin an existing scorer with scorer_id and scorer_version (scorerId, scorerVersion). Eval then writes nothing to the scorer.
  • The digest covers the function’s source text only. A changed constant it closes over, or a helper it calls, does not create a version. Put thresholds inside the function, or bump registry_name when behaviour changes.

Catalog

Each entry says what it measures, what it needs, and gives code to copy. Every example assumes the dataset shape noted under Needs.

Exact match

Measures whether the output equals the expected value, after normalising case and whitespace. Needs expected as a string.

Contains required facts

Measures the share of required phrases that appear in the output. Partial credit shows how much was missed. Needs expected as {"must_include": [...]}.

Valid JSON with required keys

Measures whether the output parses as JSON and carries every required key. Needs expected as {"keys": [...]}.
Import ScoreValue from atlanai.

Numeric within tolerance

Measures whether a numeric answer is within a relative tolerance. The tolerance lives on the case. Needs expected as a number and extra.tolerance on the record, or metadata.tolerance inline.

Pattern and refusal checks

Measures a format rule or a forbidden pattern, such as a leaked internal ID, or a refusal where an answer was required. Needs nothing beyond the output.

Only where it applies

Measures a rule that only some cases exercise. Return None or null elsewhere, so the mean is taken over the cases the rule covers. Needs a category on the case.

Tool calls and trajectories

Scoring which tools an agent called, in what order, and with what arguments is the core of agent evaluation. It has its own page: evaluate an agent.

An LLM judge in your own code

Measures a quality that code cannot decide, such as tone or faithfulness, using your own model client. Needs a model you are approved to send case data to. The hosted llm_judge scorers run over recorded sessions, not inside Eval. When a CI run needs a judge, call your model from a code scorer, parse a constrained answer, and keep the rationale:
Keep judge scorers few and deterministic: set temperature=0, constrain the answer to fixed labels, and delimit the case content as data. When you change the rubric, the source digest changes and a new scorer version is recorded. Compare runs only when they were scored by the same version.

Choosing scores that gate well

  • One behaviour per score. A score that mixes three rules cannot tell you which one regressed.
  • Partial credit where it carries information. “3 of 4 facts” is more useful than a pass or fail.
  • Abstain rather than pass. Return None when a rule does not apply. A free 1.0 hides regressions in the mean.
  • Stable names. The CI gate and trend views key on the score name. Renaming a score starts a new series.