Eval
Eval
Measure model quality on your own traffic. Compare models automatically with AutoEvals.
Model providers ship updates constantly and prompts drift. Eval gives you a repeatable way to measure model quality and to compare model options for a given task, so you can decide with data instead of intuition.
AutoEvals: compare models on your traffic
AutoEvals answer one question: is there a better model for this workload than the one you run today? A run samples your recent production traffic from a Gateway task or a traced agent, replays each sample against a set of candidate models, scores every output with LLM judges, and publishes a model recommendation with the evidence behind it.
You do not build a dataset or write a rubric. Start a run from the Compare Models tab on any task or agent, and read the results down to the individual judged sample.
Run an AutoEval
From integration to a completed run: configuration, sampling, scoring, and limits.
Interpret the results
The recommendation, the leaderboard, and the Sample Viewer.
How scoring works
Each sampled input is replayed against every candidate model, and an LLM judge scores each output against two fixed rubrics, one for quality and one for similarity to your current model. Scores aggregate into a leaderboard, and the run recommends the model that wins on quality at the lowest cost. You can open any judged sample to see the exact reasoning behind its score.
When to go further
AutoEvals tell you whether a better off-the-shelf model exists for a task you already run. When the best available model still is not good enough, AutoTrainer distills the task onto a smaller, cheaper custom model from the same traffic.
Next steps
Run an AutoEval
Configuration, sampling, scoring, and limits.
Interpret the results
Read the verdict, the leaderboard, and individual judged samples.
Compare Models get-started
A guided first run from integration to a completed comparison.
Train a custom model
Push past off-the-shelf quality with a model distilled from your traffic.