Skip to main content
A low online score is a lead for review, not an automatic ground truth label. Read the session and its trace, check whether the scorer was right, then write a case that expresses the expected behavior. Do not copy a full transcript into a dataset record when a short, redacted input reproduces the failure.

Prepare a reviewed case

For this synthetic example, a session showed an answer that omitted a required safety step. A reviewer distilled it to a question and an expected answer. The source_ref links the case to the session from which it was curated.
source_kind: "session" requires a readable session in the same authorized scope. If you cannot retain that link, omit both source_kind and source_ref and keep the review record in your approved system.

Run it on the next candidate

Pass dataset=suite.dataset.id to Eval with the candidate task and your score functions. The experiment pins the dataset snapshot. A later correction to this case creates a new record version and will affect only new experiments. Review the new case’s trace before considering the failure fixed. A correct final sentence can hide a wrong tool call or an unsafe intermediate step; add a scorer for that behavior if it matters to the release decision.