Eval records the score on the case’s result and trace, and registers the scorer as a versioned artifact so every score says which definition produced it.
Use code scorers for anything you can check deterministically: equality, structure, tool calls, keywords, budgets. Reach for an LLM judge only for what code cannot decide. For judges over production sessions, see hosted judges.
The contract
input, output, expected, and metadata, so declare only what you use and accept **_ for the rest. TypeScript passes one object with input, output, expected, metadata, tags, and id. metadata is the inline case’s metadata, or the dataset record’s extra. In TypeScript, expected is optional on every case, so type it expected?: and return null when it is missing; a required expected fails tsc --strict.
Passing scorers to Eval. Python accepts bare functions, named after the function, or {"name": ..., "scorer": fn} mappings. A lambda should always be wrapped in a mapping with a name. TypeScript accepts { name, scorer } objects only.
Return values.
A non-finite number raises. Two scorers that emit the same score name fail the run: every score name belongs to exactly one scorer.
Failures. If a scorer raises, its message is appended to that case’s
error and the other scorers still run. If the task raises, the case records the error and its scorers are skipped.
Versioning and names
Eval registers each scorer as <eval-name>-<scorer-name> in the workspace, lowercased with _ turned into -: scorer exact_match in eval support-agent is support-agent-exact-match. Score names on results keep the function’s own name, exact_match. Eval stores a digest of the function’s source. When the source changes, the next run records a new scorer version, and each score cites the version that produced it.
- Keep the
Evalname stable. Renaming the eval creates new scorer artifacts, so a baseline and a candidate with different eval names are not scored by the same scorer. The CI action also matches baselines on this name. - Share one scorer across evals by passing
registry_name(registryNamein TypeScript) in the scorer mapping. - Pin an existing scorer with
scorer_idandscorer_version(scorerId,scorerVersion).Evalthen writes nothing to the scorer. - The digest covers the function’s source text only. A changed constant it closes over, or a helper it calls, does not create a version. Put thresholds inside the function, or bump
registry_namewhen behaviour changes.
Catalog
Each entry says what it measures, what it needs, and gives code to copy. Every example assumes the dataset shape noted under Needs.Exact match
Measures whether the output equals the expected value, after normalising case and whitespace. Needsexpected as a string.
Contains required facts
Measures the share of required phrases that appear in the output. Partial credit shows how much was missed. Needsexpected as {"must_include": [...]}.
Valid JSON with required keys
Measures whether the output parses as JSON and carries every required key. Needsexpected as {"keys": [...]}.
ScoreValue from atlanai.
Numeric within tolerance
Measures whether a numeric answer is within a relative tolerance. The tolerance lives on the case. Needsexpected as a number and extra.tolerance on the record, or metadata.tolerance inline.
Pattern and refusal checks
Measures a format rule or a forbidden pattern, such as a leaked internal ID, or a refusal where an answer was required. Needs nothing beyond the output.Only where it applies
Measures a rule that only some cases exercise. ReturnNone or null elsewhere, so the mean is taken over the cases the rule covers. Needs a category on the case.
Tool calls and trajectories
Scoring which tools an agent called, in what order, and with what arguments is the core of agent evaluation. It has its own page: evaluate an agent.An LLM judge in your own code
Measures a quality that code cannot decide, such as tone or faithfulness, using your own model client. Needs a model you are approved to send case data to. The hostedllm_judge scorers run over recorded sessions, not inside Eval. When a CI run needs a judge, call your model from a code scorer, parse a constrained answer, and keep the rationale:
temperature=0, constrain the answer to fixed labels, and delimit the case content as data. When you change the rubric, the source digest changes and a new scorer version is recorded. Compare runs only when they were scored by the same version.
Choosing scores that gate well
- One behaviour per score. A score that mixes three rules cannot tell you which one regressed.
- Partial credit where it carries information. “3 of 4 facts” is more useful than a pass or fail.
- Abstain rather than pass. Return
Nonewhen a rule does not apply. A free1.0hides regressions in the mean. - Stable names. The CI gate and trend views key on the score name. Renaming a score starts a new series.