Module
Module 3 of 6Lesson 2 of 2~14 min

Label, version and protect the test set

How to get a test set labeled consistently, version it and shield it from overfitting. For a PM, this is what makes results comparable from one version to the next, and it keeps you from celebrating a rising score while the feature itself has not improved.

Lesson objective

By the end of this lesson, you will be able to have a test set labeled with a guide, check inter-annotator agreement, version it with the prompt and model, and protect it from overfitting with a holdout set.

Topics covered

  • data labeling
  • inter-annotator agreement
  • test set versioning
  • holdout set
  • prompt overfitting

Where it fits

Build the test set

Which cases do you test on, and how do you keep the set reliable over time?

Lessons in this module

  1. Build a representative test set
  2. Label, version and protect the test set (this lesson)

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.