Skip to content
Module
Module 4 of 6Lesson 1 of 3~10 min

Manage context and conversation memory

A model remembers nothing between two calls: the application decides what it sees again, and that choice weighs on quality, cost and latency. This lesson helps you choose the right memory strategy for your assistant and explain its trade-offs to your team.

Lesson objective

By the end of this lesson, you will be able to choose a memory strategy (full history, sliding window, summary, persistent memory) by weighing quality, cost and latency.

Topics covered

  • conversation memory
  • context window
  • AI assistant
  • LLM cost and latency
  • chat history

In the glossary

Full glossary

Where it fits

Give it the right knowledge

How does the assistant get the right information at the right time?

Lessons in this module

  1. Manage context and conversation memory (this lesson)
  2. Choose how the assistant gets its knowledge
  3. Specify a RAG system with your team

What you will learn in the course

This lesson is part of the course Build an AI assistant for your product

  • Identify a use case that justifies an AI assistant and write its framing brief (problem, users, allowed actions, out of scope, success criteria).
  • Design the conversational experience: entry point, tone, handling uncertainty, handoff to a human and response format.
  • Choose and justify a knowledge strategy (instructions, injected context, RAG, fine-tuning) and a conversation memory strategy.
  • Specify the assistant's tools and actions (data read, actions written, confirmation, permissions) and choose how to build it.
  • Identify the risks (injection, leaks, excessive actions, costs) and design layered guardrails that go beyond the prompt.
  • Design an evaluation plan with a reference dataset, criteria, grading methods, release thresholds and a regression rule.
  • Define production metrics (product, quality, cost, latency), alert thresholds and the loop from user feedback to the evaluation set.