AI evals (evaluating an AI feature)
Evals are systematic tests of an AI feature's outputs against defined criteria, run on a set of representative inputs. Each case pairs an input with what a good answer must contain or avoid, and a grading method scores the result: exact match, code checks, human review or another model acting as a judge. Evals turn “it seems better” into a number you can track across versions.
Why it matters for a PM
LLM outputs vary, and a prompt or model change that fixes one case can quietly break ten others. Evals are how a PM decides whether a feature is ready to ship, whether a new model is worth adopting and whether quality is drifting in production. They play the role of acceptance criteria for AI, and the PM is the natural owner of what good means for users.
Example
Before switching its email reply assistant to a cheaper model, a team runs both versions on 200 real customer emails. The cheaper model scores the same on tone but drops two points on correct policy answers, including one severe error on refunds. The team keeps the current model for that category.
Key points
- Start from success criteria tied to user impact, with a severity for each type of error.
- Build the test set from real inputs, edge cases and adversarial ones included, not only typical examples.
- Pick a grading method per criterion: code where the rule is objective, people or calibrated judges where it is not.
- Agree on the bar for shipping, and on what counts as a regression, before seeing any score.
- Keep part of the data aside so the prompt is not tuned to the test set itself.
Common mistakes
- Shipping on a vibe check of a few hand-picked examples.
- Reporting one average score that hides severe failures on a small segment.
- Never updating the test set as real usage reveals new cases.
- Evaluating offline only, and never checking quality on live traffic.
Go further with Module
The courses and lessons that cover this concept: