Start with a small offline dataset that represents the behavior you need. Compare candidates against the same cases. Then score completed sessions to find failure modes the dataset missed. Add reviewed failures back to the dataset.
Run an offline eval
Install the tracing SDK, run cases, read the summary, and resume an interrupted run.
Score production sessions
Define a judge, preview its prompt, test it, and start an asynchronous scoring run.
Compare two agent versions
Use one dataset for a baseline and a candidate so each case has a direct comparison.
Turn a failure into a test
Review a low score and add a reproducible case to the next offline run.
The objects you will see
- A dataset holds curated cases. Each record keeps its input, optional expected result, categories, and provenance.
- An experiment is one offline run over a dataset or inline cases. It pins the run configuration and records the case results.
- A scorer defines what a score means. Offline
Evalcalls your score function. Online scoring runs a storedllm_judgedefinition against session evidence. - A trace shows the task, model, and tool work behind a result. An offline result points to its case trace; an online insight points to the judged session or trace.