Skip to content
Module
AI

Context window

The context window is the maximum amount of text, counted in tokens, that a language model can take into account in a single request: instructions, conversation history, documents and its own answer included. Tokens are chunks of text; in English one token averages roughly three quarters of a word. Whatever falls outside the window is invisible to the model for that request.

Why it matters for a PM

The context window shapes many product choices: how much of a conversation an assistant remembers, whether a long contract can be analyzed in one pass, what each request costs and how long it takes. Large windows do not remove the need to select, because quality tends to drop when the relevant information is buried in a very long input.

Example

A legal assistant can accept a 300-page contract in one request, yet it misses a clause buried on page 180. The team changes the design: the contract is split by section, the assistant retrieves the sections relevant to each question and cites the page of every clause it mentions.

Key points

  • A model keeps no memory between requests: the application resends whatever it should remember.
  • Cost and latency grow with the number of tokens sent, so long histories get summarized or trimmed.
  • Details placed midway through a very long input get overlooked more often than those near its start or end.
  • Retrieval (RAG) picks the relevant passages instead of sending everything.
  • Coding agents manage their window actively: clearing it, compacting it or delegating to subagents.

Common mistakes

  • Equating a million-token window with a million tokens of reliable understanding.
  • Letting a long chat grow until early instructions are crowded out.
  • Forgetting that the answer itself consumes part of the window and of the output budget.

Go further with Module

The courses and lessons that cover this concept: