Read classification metrics
How to read a confusion matrix and pick the metric that truly reflects what a classifier's errors cost. You avoid approving a model on a flattering overall figure that hides a poorly handled critical category, and you speak the same language as the data team.
Lesson objective
By the end of this lesson, you will be able to read a confusion matrix, choose between accuracy, precision, recall and F1 based on the cost of errors, and spot when class imbalance makes a score misleading.
Topics covered
- confusion matrix
- precision and recall
- F1 score
- class imbalance
- classification metrics
Where it fits
Choose the metrics
Which metrics should you read, and what do they say about the stage to fix?
Lessons in this module
- Read classification metrics (this lesson)
- Read RAG metrics
What you will learn in the course
This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges
- Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
- Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
- Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
- Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
- Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
- Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
- Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Ship and monitor an AI feature in productionExpert · ~3 hr