Evaluate an AI feature: test sets, metrics and LLM judges
Decide on evidence that an AI feature is ready: criteria, versioned test set, grading, calibrated judge, metrics and release gates.
Decide on evidence, not impressions, that an AI feature is ready for production: success criteria, a representative and versioned test set, a grading method per criterion, a calibrated LLM judge, the right metrics and release thresholds. You leave with your feature's evaluation plan, its test set and its thresholds, without writing code.
What you will be able to do
- Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
- Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
- Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
- Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
- Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
- Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
- Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.
Prerequisites
- Have designed or managed a feature that uses an LLM (assistant, summary, classification, extraction), or be about to
- Know what a prompt, a system prompt and, for the RAG part, passage retrieval are
- Recommended, not required: Build an AI assistant for your product; Design a RAG architecture that fits your product.
- This course is not for developers looking for a tooling tutorial (SDK, code).
Syllabus
What will I build in this course, and in what order?
- Course overview3 steps
Objective · By the end of this overview, you will know what you are going to produce (your AI feature's evaluation plan, its versioned test set and its release thresholds) and in what order the five modules get you there.
What proves that an AI feature works, and against which criteria?
- Why vibes are not enough5 steps
Objective · By the end of this lesson, you will be able to explain why a few tries do not prove an AI feature's quality, and to describe what a replayable evaluation adds on every change.
Objective · By the end of this lesson, you will be able to turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, covering several dimensions and ranked by error severity.
Which cases do you test on, and how do you keep the set reliable over time?
Objective · By the end of this lesson, you will be able to compose a test set from real traffic, edge cases and adversarial cases, set its starting size and justify its split.
Objective · By the end of this lesson, you will be able to have a test set labeled with a guide, check inter-annotator agreement, version it with the prompt and model, and protect it from overfitting with a holdout set.
How do you grade each criterion reliably, and when can you trust an LLM judge?
Objective · By the end of this lesson, you will be able to choose, for each criterion, between code-based grading, human grading and an LLM judge, and to justify the choice by the nature of the criterion, cost, speed and reliability.
Objective · By the end of this lesson, you will be able to specify an LLM judge (rubric, grading format, choice of model), identify its biases and calibrate it against human grades before trusting it with evaluation at scale.
Which metrics should you read, and what do they say about the stage to fix?
- Read classification metrics5 steps
Objective · By the end of this lesson, you will be able to read a confusion matrix, choose between accuracy, precision, recall and F1 based on the cost of errors, and spot when class imbalance makes a score misleading.
- Read RAG metrics5 steps
Objective · By the end of this lesson, you will be able to choose the metrics for a RAG-based feature (context recall and precision, faithfulness, answer correctness, citation accuracy), know which ones require a reference, and infer which stage to fix.
When is a version ready, and how do you keep a change from degrading it?
Objective · By the end of this lesson, you will be able to define release thresholds set before seeing results, a regression rule for every change, and the role of offline and online evaluations in the decision.
Objective · By the end of this lesson, you will be able to assemble your AI feature's evaluation plan, with its versioned test set and release thresholds, and plan its rollout at 7 and 30 days.