Manage context and conversation memory
A model remembers nothing between two calls: the application decides what it sees again, and that choice weighs on quality, cost and latency. This lesson helps you choose the right memory strategy for your assistant and explain its trade-offs to your team.
Lesson objective
By the end of this lesson, you will be able to choose a memory strategy (full history, sliding window, summary, persistent memory) by weighing quality, cost and latency.
Topics covered
- conversation memory
- context window
- AI assistant
- LLM cost and latency
- chat history
In the glossary
Where it fits
Give it the right knowledge
How does the assistant get the right information at the right time?
Lessons in this module
- Manage context and conversation memory (this lesson)
- Choose how the assistant gets its knowledge
- Specify a RAG system with your team
What you will learn in the course
This lesson is part of the course Build an AI assistant for your product
- Identify a use case that justifies an AI assistant and write its framing brief (problem, users, allowed actions, out of scope, success criteria).
- Design the conversational experience: entry point, tone, handling uncertainty, handoff to a human and response format.
- Choose and justify a knowledge strategy (instructions, injected context, RAG, fine-tuning) and a conversation memory strategy.
- Specify the assistant's tools and actions (data read, actions written, confirmation, permissions) and choose how to build it.
- Identify the risks (injection, leaks, excessive actions, costs) and design layered guardrails that go beyond the prompt.
- Design an evaluation plan with a reference dataset, criteria, grading methods, release thresholds and a regression rule.
- Define production metrics (product, quality, cost, latency), alert thresholds and the loop from user feedback to the evaluation set.
Related courses
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr
- Ship and monitor an AI feature in productionExpert · ~3 hr