Recognise AI-specific incidents
Lesson 1 of the module "Incidents, fallbacks and model changes" in the course "Ship and monitor an AI feature in production".
Lesson objective
By the end of this lesson, you will be able to recognise the incidents specific to an AI feature (quality drift, provider unavailability, model deprecation, call loop, cost spike, misuse), classify them by severity and write the runbook that says what to do for each.
Where it fits
Incidents, fallbacks and model changes
What do you do when quality, the provider or the model shifts under your feet?
Lessons in this module
- Recognise AI-specific incidents (this lesson)
- Plan the fallbacks
- Change models without regressions
What you will learn in the course
This lesson is part of the course Ship and monitor an AI feature in production
- Design the staged launch of an AI feature (feature flag, shadow, canary, percentages) with quantified promotion criteria and a tested kill switch.
- Define what you log and trace for each model call, what you mask or exclude (personal data), how long you keep it, and with which tool.
- Set up quality signals in production (user feedback, sampled human review, online evaluations, drift) and connect them to the evaluation set.
- Manage cost per request and per useful outcome, and perceived latency (budgets, alerts, caching, streaming, response length, model choice).
- Diagnose AI-specific incidents (drift, provider outage, deprecation, call loop, cost spike) and plan the right fallbacks.
- Run a model or prompt change with regression evaluations, shadow comparison, staged rollout and a rollback plan.
- Build the quality, cost and latency dashboard of an AI feature and report to stakeholders with decisions to make.
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr