The verdict
The report leads with the recommendation: either Recommended model with a specific alternative, or Keep current model.
- Eval score vs current: how much better or worse the recommended model scored than your current one.
- Monthly savings: what switching would save, projected from the real traffic volume in the sampled window.
- View judge notes: the analysis agent’s reasoning for the pick.
- Switch to this model: a ready-made prompt you can paste into a coding agent to migrate your task or agent off the current model.
Key findings and run summary
Key findings lists the most important observations from the judged samples, marked as positive, caution, or negative. Below it, the Run summary is a written analysis of the run with citations that link directly into the underlying requests and traces, so every claim can be checked against real traffic. Findings often go beyond model choice: the analysis agent also records configuration problems and failure modes it noticed in your traffic.Quality versus cost
The scatter chart plots every candidate by score against cost. Up and to the left is better: higher quality at lower cost. The current and recommended models are labeled. A second tab, Similarity vs. cost, shows the same chart with the similarity score on the y-axis.
The leaderboard
The Eval scores table ranks every candidate against your current model.
Use the Rubric selector to rank by quality, similarity, or both. Click a row to expand the judge’s verdict on that model: a narrative summary of how it behaved, and a View judged samples link into the raw results.
Models that could not be scored appear as muted rows with the skip reason.
Reading quality against similarity
The two scores answer different questions, and the combination matters:- High quality, high similarity: behaves like your current model and scores well. The safest switch.
- High quality, low similarity: scores well but behaves differently (different tool choices, different response style). Test it before routing traffic to it.
- Low quality: not a candidate, regardless of similarity.
The Sample Viewer
Every score in the report is backed by judged samples you can read. The Inspect outputs across all models card at the bottom of the report (and the Sample Viewer button on the leaderboard) opens the eval viewer for the run. The eval viewer shows:- Score distribution per model, so you can see spread and outliers rather than just averages. Click a model to filter.
- Comparison grid: one row per sample, with the original input and output pinned on the left and one column per model. Click any cell to open the full result, including the judge’s scored assessment.
- Sort controls: sort samples by any model’s score, highest or lowest first. Sorting by your current model, lowest first, surfaces the traffic your production model handles worst.
When a run ends without results
Not every run produces a verdict. The report states what happened and why:
Each state offers Run new comparison to try again, and the limit case links to your plan and usage. The daily eval limit resets at midnight PT.
Next steps
Run an Auto Eval
Configuration, sampling, scoring, and limits.
Write your own rubric
Measure a quality dimension the fixed rubrics do not cover.
Run a Direct Eval
Compare models on a curated dataset with your own rubric.
Train your own model
Push past off-the-shelf quality with a model distilled from your traffic.