Skip to content
Module
AI

RAG (retrieval-augmented generation)

Retrieval-augmented generation (RAG) is an architecture where an application first searches a document collection for the excerpts that best match a question, then sends those passages to a large language model along with the question. The model writes its answer from that supplied context rather than from memory alone. The term comes from a 2020 research paper by Lewis and colleagues at Facebook AI Research.

Why it matters for a PM

Most AI assistants built on company knowledge (help centers, contracts, internal wikis, product catalogs) are RAG systems. As a PM you decide what the assistant may read, how fresh the content must be, who is allowed to see which documents, and what the product shows when nothing relevant is found. Those are product decisions, and they often weigh more on answer quality than the choice of model.

Example

A support assistant for an accounting app answers “Can I export invoices to my bank format?” by first retrieving the three most relevant help articles, then generating a short answer that links to them. If no article covers the bank in question, it says so and offers to open a ticket instead of guessing.

Key points

  • RAG has two halves: retrieval (finding the right passages) and generation (writing an answer from them). Most quality problems start in retrieval.
  • It suits knowledge that changes often or is private: updating an index is faster and cheaper than retraining a model.
  • Documents are split into chunks and indexed, usually with embeddings for semantic search, often combined with keyword search.
  • Good RAG products cite their sources and say plainly when the documents do not contain the answer.
  • Access rights are applied during retrieval, so a user never gets an answer built from a document they cannot open.

Common mistakes

  • Treating RAG as a cure for hallucinations: the model can still misread or ignore the passages it receives.
  • Indexing everything without owners or freshness rules, so outdated pages keep surfacing in answers.
  • Judging quality from a handful of demos instead of a test set of real questions with expected answers.
  • Choosing fine-tuning to add knowledge, when fine-tuning mostly changes style and behavior.

Go further with Module

The courses and lessons that cover this concept: