Module
Module 4 of 6Lesson 2 of 2~19 min

Red teaming, abuse monitoring and incident response

Lesson 2 of the module "Guardrails and security operations" in the course "Secure an AI product: prompt injection, guardrails and the AI Act".

Lesson objective

By the end of this lesson, you will be able to plan a red teaming campaign and the adversarial test set that comes out of it, choose the abuse monitoring signals in production, and write the incident runbook of an AI feature, kill switch and notifications included.

Where it fits

Guardrails and security operations

Which guardrails do you specify, how do you know they hold, and what do you do on incident day?

Lessons in this module

  1. Specify layered, testable and costed guardrails
  2. Red teaming, abuse monitoring and incident response (this lesson)

What you will learn in the course

This lesson is part of the course Secure an AI product: prompt injection, guardrails and the AI Act

  • Map the attack surface of an AI product and its risks with the OWASP Top 10 for LLM Applications 2025, in a prioritized risk register.
  • Analyze a product's direct and indirect prompt injection paths, system prompt leakage and data exfiltration (lethal trifecta).
  • Specify layered guardrails (input, output, rights, human approval, isolation, limits) with testable criteria and their cost.
  • Plan red teaming, abuse monitoring and incident response for an AI feature.
  • Qualify a product under the AI Act (role, risk level, prohibited practices, transparency, timeline) and connect this analysis with the GDPR.