Module
Module 6 of 6Lesson 2 of 3~16 min

Cost the RAG system and choose how to build it

The cost of RAG goes well beyond the model bill, and the way it is built commits the product for a long time. The lesson helps a PM identify cost items, compare hosted tools, managed knowledge bases and custom builds, and choose on criteria that actually matter for their product.

Lesson objective

By the end of this lesson, you will be able to list the cost items of a RAG system and choose between a model provider's hosted tool, a managed knowledge base, vector search in an existing database and a custom pipeline.

Topics covered

  • RAG cost
  • build vs buy
  • managed knowledge base
  • vector search

Where it fits

Test, cost and document

How do I prove my RAG works, what does it cost, and how do I document it?

Lessons in this module

  1. Build the 30-question test set and diagnose failures
  2. Cost the RAG system and choose how to build it (this lesson)
  3. Document your product's RAG architecture

What you will learn in the course

This lesson is part of the course Design a RAG architecture that fits your product

  • Choose, for an information need, between long context, data injected by the application, RAG and fine-tuning, and justify the choice by volume, update frequency, access rights and cost.
  • Specify the source inventory, exclusions, the metadata to capture and the document chunking strategy.
  • Design retrieval: semantic, keyword or hybrid search, reranking, contextual retrieval, number of passages and relevance threshold.
  • Specify the answer grounded in sources, how citations are displayed, how conflicting sources are handled and what happens when nothing is found.
  • Specify access filtering before the model reads any passage, and index freshness (resync, deletions, versions).
  • Build a set of 30 test questions that covers RAG failure modes and diagnose which stage caused a wrong answer.
  • Estimate the cost items of a RAG system and choose between a hosted tool, a managed knowledge base, a vector database and a custom build.