Module
Module 4 of 6Lesson 1 of 2~22 min

Improve retrieval with hybrid search, reranking and context

When the assistant fails to find the right passage, several techniques can help, but each one adds latency and cost. You learn what hybrid search, reranking, contextual retrieval and query rewriting each fix, in what order to test them, and how to judge their real effect on your product.

Lesson objective

By the end of this lesson, you will be able to choose which retrieval improvements to test (hybrid search, reranking, contextual retrieval, query rewriting), in what order, and to judge their effect on quality, latency and cost.

Topics covered

  • hybrid search
  • reranking
  • contextual retrieval
  • BM25
  • query rewriting

Where it fits

Retrieve and answer

How do you find the right passages, and how should the assistant answer from them?

Lessons in this module

  1. Improve retrieval with hybrid search, reranking and context (this lesson)
  2. Specify the cited answer and "not found"

What you will learn in the course

This lesson is part of the course Design a RAG architecture that fits your product

  • Choose, for an information need, between long context, data injected by the application, RAG and fine-tuning, and justify the choice by volume, update frequency, access rights and cost.
  • Specify the source inventory, exclusions, the metadata to capture and the document chunking strategy.
  • Design retrieval: semantic, keyword or hybrid search, reranking, contextual retrieval, number of passages and relevance threshold.
  • Specify the answer grounded in sources, how citations are displayed, how conflicting sources are handled and what happens when nothing is found.
  • Specify access filtering before the model reads any passage, and index freshness (resync, deletions, versions).
  • Build a set of 30 test questions that covers RAG failure modes and diagnose which stage caused a wrong answer.
  • Estimate the cost items of a RAG system and choose between a hosted tool, a managed knowledge base, a vector database and a custom build.