Skip to content
Module
AI

LLM as a judge

LLM as a judge is an evaluation method in which a language model scores or compares the outputs of an AI feature against a rubric. It covers qualities that code cannot check, such as relevance, tone, faithfulness to a source or completeness. The judge receives the input, the output and sometimes a reference answer, then returns a grade with a justification.

Why it matters for a PM

Human review does not scale to hundreds of cases on every prompt change. A judge model makes frequent evaluation affordable, but only if its grades agree with those of people who know what good looks like. The PM owns the rubric and decides how much a release may depend on an automated grader.

Example

A team checks whether its meeting summaries capture every decision. A judge model receives the transcript and the summary, and answers one question: is any decision from the transcript missing, yes or no, with the quote. Before relying on it, the team compares its verdicts with a PM's on 50 meetings.

Key points

  • A judge is a model too: it needs a precise rubric, examples and one narrow question per criterion.
  • Calibrate it against human labels on a sample and measure their agreement before trusting it.
  • Known biases include favoring longer answers, the option shown first, and outputs from its own model family.
  • Pass or fail verdicts and short scales are usually more consistent than scores from 1 to 10.

Common mistakes

  • Asking a judge “is this answer good?” without criteria.
  • Grading an answer with the same model and prompt that produced it.
  • Never rechecking the judge after changing its model or its rubric.
  • Replacing every human review, including on high-severity cases.

Go further with Module

The courses and lessons that cover this concept: