Module
Module 6 of 6Lesson 1 of 2~18 min

Set release thresholds and the regression rule

How to set the conditions a new version of an AI feature must meet before it ships, and what to rerun after every change. This framework spares you last-minute debates during a model upgrade or a prompt edit, and it clarifies the role of offline tests versus observation in production.

Lesson objective

By the end of this lesson, you will be able to define release thresholds set before seeing results, a regression rule for every change, and the role of offline and online evaluations in the decision.

Topics covered

  • release thresholds
  • AI regression testing
  • offline evaluation
  • online evaluation
  • model upgrade

Where it fits

Decide on release

When is a version ready, and how do you keep a change from degrading it?

Lessons in this module

  1. Set release thresholds and the regression rule (this lesson)
  2. Assemble your feature's evaluation plan

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.