Build the 30-question test set and diagnose failures
Without a test set, there is no way to tell whether a RAG system is improving or to compare two solutions. This lesson shows a PM how to build test questions that cover the different ways RAG fails, then how to trace a bad answer back to the pipeline stage responsible for it.
Lesson objective
By the end of this lesson, you will be able to build a set of 30 test questions that covers RAG failure modes, and to diagnose which stage caused a wrong answer (retrieval, ranking, freshness, generation).
Topics covered
- RAG evaluation
- test set
- failure modes
- answer faithfulness
- context recall
Where it fits
Test, cost and document
How do I prove my RAG works, what does it cost, and how do I document it?
Lessons in this module
- Build the 30-question test set and diagnose failures (this lesson)
- Cost the RAG system and choose how to build it
- Document your product's RAG architecture
What you will learn in the course
This lesson is part of the course Design a RAG architecture that fits your product
- Choose, for an information need, between long context, data injected by the application, RAG and fine-tuning, and justify the choice by volume, update frequency, access rights and cost.
- Specify the source inventory, exclusions, the metadata to capture and the document chunking strategy.
- Design retrieval: semantic, keyword or hybrid search, reranking, contextual retrieval, number of passages and relevance threshold.
- Specify the answer grounded in sources, how citations are displayed, how conflicting sources are handled and what happens when nothing is found.
- Specify access filtering before the model reads any passage, and index freshness (resync, deletions, versions).
- Build a set of 30 test questions that covers RAG failure modes and diagnose which stage caused a wrong answer.
- Estimate the cost items of a RAG system and choose between a hosted tool, a managed knowledge base, a vector database and a custom build.
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr
- Ship and monitor an AI feature in productionExpert · ~3 hr