Skip to content
Module
AI

Embeddings and semantic search

An embedding is a list of numbers, a vector, that represents the meaning of a piece of text or an image, so that items with similar meanings get vectors close to each other. Semantic search converts the query into an embedding and retrieves the stored passages whose vectors are nearest, often measured by cosine similarity. It finds a page about how to “end your plan” when the user typed “cancel my subscription”.

Why it matters for a PM

Embeddings sit behind RAG assistants, similar-item recommendations, duplicate ticket detection and search bars that understand intent. A PM can skip the math, yet should understand what drives relevance: how documents are split, which embedding model is used, how results are filtered and ranked, and how retrieval quality is measured on real queries.

Example

In an internal wiki, an employee searches for “parental leave”. Keyword search finds nothing because the page is titled “Family policy”. Semantic search returns it first, since the two phrases have close embeddings, and a filter limits results to pages from the employee's own country.

Key points

  • Vectors live in a vector database or index, usually alongside metadata that allows filtering.
  • Semantic search can miss exact terms such as product codes or names; hybrid search adds keyword matching such as BM25.
  • A reranking step can reorder the top results with a more precise model.
  • Changing the embedding model means computing the vectors of the whole corpus again.

Common mistakes

  • Returning the nearest results even when none is relevant, instead of applying a threshold.
  • Splitting documents into chunks too small to carry meaning, or too large to be specific.
  • Assuming access rights are respected because the search is fast and accurate.

Go further with Module

The courses and lessons that cover this concept: