Module
Module 5 of 6Lesson 1 of 2~16 min

Read classification metrics

How to read a confusion matrix and pick the metric that truly reflects what a classifier's errors cost. You avoid approving a model on a flattering overall figure that hides a poorly handled critical category, and you speak the same language as the data team.

Lesson objective

By the end of this lesson, you will be able to read a confusion matrix, choose between accuracy, precision, recall and F1 based on the cost of errors, and spot when class imbalance makes a score misleading.

Topics covered

  • confusion matrix
  • precision and recall
  • F1 score
  • class imbalance
  • classification metrics

Where it fits

Choose the metrics

Which metrics should you read, and what do they say about the stage to fix?

Lessons in this module

  1. Read classification metrics (this lesson)
  2. Read RAG metrics

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.