Module
Module 5 of 6Lesson 2 of 2~18 min

Read RAG metrics

Which metrics to track for a feature built on RAG, and how they point to the stage that needs fixing. Faced with a wrong answer, you can tell a retrieval problem from a generation problem together with the engineering team, instead of tweaking the prompt at random.

Lesson objective

By the end of this lesson, you will be able to choose the metrics for a RAG-based feature (context recall and precision, faithfulness, answer correctness, citation accuracy), know which ones require a reference, and infer which stage to fix.

Topics covered

  • RAG metrics
  • faithfulness
  • context recall
  • context precision
  • citation accuracy

Where it fits

Choose the metrics

Which metrics should you read, and what do they say about the stage to fix?

Lessons in this module

  1. Read classification metrics
  2. Read RAG metrics (this lesson)

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.