Skip to content
Module
Module 6 of 6Lesson 1 of 4~15 min

Design the evaluation plan

Testing a few questions by hand doesn't tell you whether your assistant is ready, or whether it regresses after the next change. This lesson teaches you to design a repeatable evaluation plan, with clear criteria, a way to grade each one and release rules set in advance.

Lesson objective

By the end of this lesson, you will be able to design an evaluation plan: reference dataset composition, criteria, a grading method per criterion, release thresholds and a regression rule.

Topics covered

  • AI assistant evaluation
  • LLM evals
  • golden dataset
  • LLM as a judge
  • release thresholds

In the glossary

Full glossary

Where it fits

Evaluate, ship, monitor

How do you know the assistant is ready, launch it safely and improve it?

Lessons in this module

  1. Design the evaluation plan (this lesson)
  2. Choose how to build it
  3. Launch and run in production
  4. Assemble the design dossier

What you will learn in the course

This lesson is part of the course Build an AI assistant for your product

  • Identify a use case that justifies an AI assistant and write its framing brief (problem, users, allowed actions, out of scope, success criteria).
  • Design the conversational experience: entry point, tone, handling uncertainty, handoff to a human and response format.
  • Choose and justify a knowledge strategy (instructions, injected context, RAG, fine-tuning) and a conversation memory strategy.
  • Specify the assistant's tools and actions (data read, actions written, confirmation, permissions) and choose how to build it.
  • Identify the risks (injection, leaks, excessive actions, costs) and design layered guardrails that go beyond the prompt.
  • Design an evaluation plan with a reference dataset, criteria, grading methods, release thresholds and a regression rule.
  • Define production metrics (product, quality, cost, latency), alert thresholds and the loop from user feedback to the evaluation set.