Build a representative test set
How to build a test set that looks like your users' real requests, including rare cases and attempts to misuse the feature. It is the foundation of every decision about an AI feature: a poorly built set yields reassuring scores on cases that don't matter, and its size drives the effort you negotiate with the data team.
Lesson objective
By the end of this lesson, you will be able to compose a test set from real traffic, edge cases and adversarial cases, set its starting size and justify its split.
Topics covered
- AI test set
- golden dataset
- stratified sampling
- edge cases
- adversarial cases
- synthetic data
Where it fits
Build the test set
Which cases do you test on, and how do you keep the set reliable over time?
Lessons in this module
- Build a representative test set (this lesson)
- Label, version and protect the test set
What you will learn in the course
This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges
- Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
- Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
- Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
- Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
- Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
- Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
- Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Ship and monitor an AI feature in productionExpert · ~3 hr