Module
Module 1 of 6Lesson 1 of 1~2 min

Course overview

This course on LLM evals helps you move from a good feeling about an AI feature to a launch decision you can defend. Working on your own product, you build the pieces that make that decision solid and that your engineers can rerun with every new version.

Lesson objective

By the end of this overview, you will know what you are going to produce (your AI feature's evaluation plan, its versioned test set and its release thresholds) and in what order the five modules get you there.

Topics covered

  • LLM evals
  • AI feature evaluation
  • LLM as a judge
  • test set
  • AI release decision

Where it fits

Overview

What will I build in this course, and in what order?

Lessons in this module

  1. Course overview (this lesson)

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.