Check out the newest way to compare different models for a task/agent harness: AutoEvals

Eval

Eval

Measure model quality on your own traffic. Compare models automatically with AutoEvals.

Model providers ship updates constantly and prompts drift. Eval gives you a repeatable way to measure model quality and to compare model options for a given task, so you can decide with data instead of intuition.

AutoEvals: compare models on your traffic

AutoEvals answer one question: is there a better model for this workload than the one you run today? A run samples your recent production traffic from a Gateway task or a traced agent, replays each sample against a set of candidate models, scores every output with LLM judges, and publishes a model recommendation with the evidence behind it.

You do not build a dataset or write a rubric. Start a run from the Compare Models tab on any task or agent, and read the results down to the individual judged sample.

How scoring works

Each sampled input is replayed against every candidate model, and an LLM judge scores each output against two fixed rubrics, one for quality and one for similarity to your current model. Scores aggregate into a leaderboard, and the run recommends the model that wins on quality at the lowest cost. You can open any judged sample to see the exact reasoning behind its score.

When to go further

AutoEvals tell you whether a better off-the-shelf model exists for a task you already run. When the best available model still is not good enough, AutoTrainer distills the task onto a smaller, cheaper custom model from the same traffic.

Next steps

On this page