Design the evaluation plan
Testing a few questions by hand doesn't tell you whether your assistant is ready, or whether it regresses after the next change. This lesson teaches you to design a repeatable evaluation plan, with clear criteria, a way to grade each one and release rules set in advance.
Lesson objective
By the end of this lesson, you will be able to design an evaluation plan: reference dataset composition, criteria, a grading method per criterion, release thresholds and a regression rule.
Topics covered
- AI assistant evaluation
- LLM evals
- golden dataset
- LLM as a judge
- release thresholds
In the glossary
Where it fits
Evaluate, ship, monitor
How do you know the assistant is ready, launch it safely and improve it?
Lessons in this module
- Design the evaluation plan (this lesson)
- Choose how to build it
- Launch and run in production
- Assemble the design dossier
What you will learn in the course
This lesson is part of the course Build an AI assistant for your product
- Identify a use case that justifies an AI assistant and write its framing brief (problem, users, allowed actions, out of scope, success criteria).
- Design the conversational experience: entry point, tone, handling uncertainty, handoff to a human and response format.
- Choose and justify a knowledge strategy (instructions, injected context, RAG, fine-tuning) and a conversation memory strategy.
- Specify the assistant's tools and actions (data read, actions written, confirmation, permissions) and choose how to build it.
- Identify the risks (injection, leaks, excessive actions, costs) and design layered guardrails that go beyond the prompt.
- Design an evaluation plan with a reference dataset, criteria, grading methods, release thresholds and a regression rule.
- Define production metrics (product, quality, cost, latency), alert thresholds and the loop from user feedback to the evaluation set.
Related courses
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr
- Ship and monitor an AI feature in productionExpert · ~3 hr