Module
Back to courses
AI & ProductAdvanced

Design a RAG architecture that fits your product

Decide whether you need RAG, then spec each stage: sources, chunking, retrieval, access rights, freshness, cited answers and cost, with 30 test questions.

61 steps~3 hr 30 minLevel: Advanced · Regular practice: you already use the tools or work on this topic.

Decide whether your product needs RAG, then specify each stage with your team: sources, chunking, retrieval, access rights, freshness, cited answers and cost. You leave with your product's documented RAG architecture and a set of 30 test questions, without writing code.

What you will be able to do

  • Choose, for an information need, between long context, data injected by the application, RAG and fine-tuning, and justify the choice by volume, update frequency, access rights and cost.
  • Specify the source inventory, exclusions, the metadata to capture and the document chunking strategy.
  • Design retrieval: semantic, keyword or hybrid search, reranking, contextual retrieval, number of passages and relevance threshold.
  • Specify the answer grounded in sources, how citations are displayed, how conflicting sources are handled and what happens when nothing is found.
  • Specify access filtering before the model reads any passage, and index freshness (resync, deletions, versions).
  • Build a set of 30 test questions that covers RAG failure modes and diagnose which stage caused a wrong answer.
  • Estimate the cost items of a RAG system and choose between a hosted tool, a managed knowledge base, a vector database and a custom build.

Prerequisites

  • Know what an LLM, a prompt and a context window are, well enough to explain them to a colleague
  • Have written specs or user stories with an engineering team
  • Recommended, not required: Build an AI assistant for your product (end-to-end view of an assistant, including a first lesson on RAG).
  • This course is not for developers looking for an implementation tutorial (SDK, code).

Syllabus

What will I design in this course, and in what order?

  1. Objective · By the end of this overview, you will know what you are going to produce (your product's documented RAG architecture and a set of 30 test questions) and in what order the five modules get you there.

Does my need justify RAG, and which sources should go into it?

  1. Objective · By the end of this lesson, you will be able to choose, for each information need of your product, between long context, data injected by the application, RAG and fine-tuning, and to justify the choice by volume, update frequency, access rights and cost.

  2. Objective · By the end of this lesson, you will be able to build a RAG source inventory from real user questions, decide what is excluded, and list the metadata to capture at ingestion.

How do documents become passages that retrieval can find?

  1. Objective · By the end of this lesson, you will be able to choose a chunking strategy suited to your documents' structure and to the expected answer unit, and to phrase it as requirements the team can test.

  2. Objective · By the end of this lesson, you will be able to explain how semantic search retrieves passages, list the product decisions it involves (model, number of passages, relevance threshold) and anticipate its weak spots.

How do you find the right passages, and how should the assistant answer from them?

  1. Objective · By the end of this lesson, you will be able to choose which retrieval improvements to test (hybrid search, reranking, contextual retrieval, query rewriting), in what order, and to judge their effect on quality, latency and cost.

  2. Objective · By the end of this lesson, you will be able to specify how the assistant answers from the passages: grounding, citation format, the rule for conflicting sources and the behavior when nothing relevant is found.

How do you guarantee that each user only reads what they are allowed to see, in its current version?

  1. Objective · By the end of this lesson, you will be able to specify a RAG system's access filtering (where rights come from, where the filter applies, how rights stay in sync) and to reject an architecture whose security relies on the prompt.

  2. Objective · By the end of this lesson, you will be able to set a freshness target per source and specify the mechanisms that meet it: resync, deletion propagation, versions and an alert on failure.

How do I prove my RAG works, what does it cost, and how do I document it?

  1. Objective · By the end of this lesson, you will be able to build a set of 30 test questions that covers RAG failure modes, and to diagnose which stage caused a wrong answer (retrieval, ranking, freshness, generation).

  2. Objective · By the end of this lesson, you will be able to list the cost items of a RAG system and choose between a model provider's hosted tool, a managed knowledge base, vector search in an existing database and a custom pipeline.

  3. Objective · By the end of this lesson, you will be able to assemble your product's documented RAG architecture in seven sections, with its 30-question test set, and plan its rollout at 7 and 30 days.