Module
Back to courses
AI & ProductExpert

Ship and monitor an AI feature in production

Launch an AI feature in stages, then know at any time what it costs, how long it takes and whether it answers well.

81 steps~3 hrLevel: Expert · Advanced level: for deeper mastery and specialization.

Launch an AI feature in stages, then know at any moment what it costs, how long it takes and whether it answers well. You learn to roll out progressively, to log without exposing personal data, to track quality, cost and latency, to respond to AI-specific incidents and to change models without regressions. You leave with the dashboard and the model change procedure for your own feature.

What you will be able to do

  • Design the staged launch of an AI feature (feature flag, shadow, canary, percentages) with quantified promotion criteria and a tested kill switch.
  • Define what you log and trace for each model call, what you mask or exclude (personal data), how long you keep it, and with which tool.
  • Set up quality signals in production (user feedback, sampled human review, online evaluations, drift) and connect them to the evaluation set.
  • Manage cost per request and per useful outcome, and perceived latency (budgets, alerts, caching, streaming, response length, model choice).
  • Diagnose AI-specific incidents (drift, provider outage, deprecation, call loop, cost spike) and plan the right fallbacks.
  • Run a model or prompt change with regression evaluations, shadow comparison, staged rollout and a rollback plan.
  • Build the quality, cost and latency dashboard of an AI feature and report to stakeholders with decisions to make.

Prerequisites

  • Have an AI feature in design, in testing or already in production (assistant, summary, classification, extraction…), or a realistic case in mind
  • Know what an evaluation set and a release threshold are. Recommended course beforehand, not required: “Evaluate an AI feature: test sets, metrics and LLM judges”.
  • Know what a token, a prompt and a model API call are, at the "I know what it is" level
  • This course is not for engineers looking to configure a specific tool: it gives the PM what they need to decide, specify and steer, without writing code.

Syllabus

What will I set up in this course, and in what order?

  1. Objective · By the end of this introduction, you will know what you are going to produce (a quality, cost, latency dashboard and a model change procedure) and in what order the four modules get you there.

How do you expose an AI feature to real users in stages, while being able to reconstruct every response?

  1. Objective · By the end of this lesson, you will be able to write the launch plan for an AI feature in stages (shadow, internal, canary, percentages), with quantified promotion criteria, an observation period and a kill switch tested before the first user.

  2. Objective · By the end of this lesson, you will be able to specify what a trace must contain for each model call, and to write a logging policy that states which fields to keep, mask or exclude, for what purpose and for how long.

How do you know, every week, whether the feature answers well, what it costs and whether it is fast enough?

  1. Objective · By the end of this lesson, you will be able to combine three quality signals in production (user feedback, sampled human review, online evaluations), spot input drift and turn failures into new cases in the evaluation set.

  2. Objective · By the end of this lesson, you will be able to calculate the cost per request and the cost per useful outcome of an AI feature, identify its main reduction levers (length, caching, model, unnecessary steps) and set budget alerts.

  3. Objective · By the end of this lesson, you will be able to choose the right latency measures (time to first token, total time, percentiles) for the intended experience, and to propose the right levers (streaming, length, model, steps, deferred processing).

What do you do when quality, the provider or the model shifts under your feet?

  1. Objective · By the end of this lesson, you will be able to recognise the incidents specific to an AI feature (quality drift, provider unavailability, model deprecation, call loop, cost spike, misuse), classify them by severity and write the runbook that says what to do for each.

  2. Objective · By the end of this lesson, you will be able to specify the fallback chain of an AI feature (bounded retries, maximum wait, backup model, degraded mode, handoff to a human) and choose, for each type of failure, the fallback that stays useful for the user.

  3. Objective · By the end of this lesson, you will be able to write the model change procedure for your AI feature: triggers, criterion-by-criterion regression evaluation, shadow comparison, staged rollout, rollback criteria and a timeline aligned with deprecation dates.

Which tools do you use to track all this, and how do you report on it to get decisions?

  1. Objective · By the end of this lesson, you will be able to tell apart the main types of observability tools for AI (dedicated tool via SDK, gateway, general-purpose monitoring platform) and choose with explicit criteria, including data protection.

  2. Objective · By the end of this lesson, you will be able to write the one-page monthly report of an AI feature, which links quality, cost and latency to the product goal, explains the gaps and ends with decisions to make.

  3. Objective · By the end of this workshop, you will have produced, for your own AI feature, a quality, cost, latency dashboard (eight metrics at most, three quantified alerts, one owner per metric) and a model change procedure, with the kit and the action plan to put them into service.