Skip to content
Module
Module 5 of 6Lesson 2 of 2~14 min

Design layered guardrails

Data leaks, off-topic answers, unauthorized actions: the risks of an embedded assistant aren't solved with one sentence in the prompt. This lesson helps you identify the threats that matter for your product and spread protections across several layers, from authorization to monitoring.

Lesson objective

By the end of this lesson, you will be able to identify the main risks of an embedded assistant and place each guardrail in the right layer (authorization, input, model, tools, output, monitoring).

Topics covered

  • AI guardrails
  • AI assistant security
  • threat model
  • prompt injection
  • responsible AI

In the glossary

Full glossary

Where it fits

Actions and guardrails

What can the assistant do, and how do you stop it from doing what it shouldn't?

Lessons in this module

  1. Specify the assistant's tools and actions
  2. Design layered guardrails (this lesson)

What you will learn in the course

This lesson is part of the course Build an AI assistant for your product

  • Identify a use case that justifies an AI assistant and write its framing brief (problem, users, allowed actions, out of scope, success criteria).
  • Design the conversational experience: entry point, tone, handling uncertainty, handoff to a human and response format.
  • Choose and justify a knowledge strategy (instructions, injected context, RAG, fine-tuning) and a conversation memory strategy.
  • Specify the assistant's tools and actions (data read, actions written, confirmation, permissions) and choose how to build it.
  • Identify the risks (injection, leaks, excessive actions, costs) and design layered guardrails that go beyond the prompt.
  • Design an evaluation plan with a reference dataset, criteria, grading methods, release thresholds and a regression rule.
  • Define production metrics (product, quality, cost, latency), alert thresholds and the loop from user feedback to the evaluation set.