Module
Module 4 of 6Lesson 2 of 2~24 min

Design and calibrate an LLM judge

How to specify a model judge, spot its biases and check that it grades like your experts before trusting it with thousands of answers. A poorly tuned judge produces flattering scores that can get a faulty version approved; knowing how to audit it lets you ask the right questions of the team that built it.

Lesson objective

By the end of this lesson, you will be able to specify an LLM judge (rubric, grading format, choice of model), identify its biases and calibrate it against human grades before trusting it with evaluation at scale.

Topics covered

  • LLM as a judge
  • model judge
  • judge bias
  • judge calibration
  • evaluation rubric

Where it fits

Grade the answers

How do you grade each criterion reliably, and when can you trust an LLM judge?

Lessons in this module

  1. Choose a grading method per criterion
  2. Design and calibrate an LLM judge (this lesson)

What you will learn in the course

This lesson is part of the course Evaluate an AI feature: test sets, metrics and LLM judges

  • Turn an AI feature's goal into specific, measurable, achievable and relevant success criteria, taking error severity into account.
  • Build a representative test set from real traffic, edge cases and adversarial cases, and justify its composition.
  • Label the test set with a guide, measure inter-annotator agreement, version it and protect it from overfitting.
  • Choose, for each criterion, a grading method (code, human, LLM judge) and justify the trade-off between cost, speed and reliability.
  • Design an LLM judge (rubric, format, different model), identify its biases and calibrate it against human grades.
  • Choose and interpret the right metrics for a classification and for a RAG system, and infer which stage to fix.
  • Define the regression rule, how offline and online evaluations fit together, and release thresholds set before seeing results.