Module
Module 3 of 5Lesson 2 of 3~18 min

Manage cost per request and per outcome

Lesson 2 of the module "Measure quality, cost and latency" in the course "Ship and monitor an AI feature in production".

Lesson objective

By the end of this lesson, you will be able to calculate the cost per request and the cost per useful outcome of an AI feature, identify its main reduction levers (length, caching, model, unnecessary steps) and set budget alerts.

Where it fits

Measure quality, cost and latency

How do you know, every week, whether the feature answers well, what it costs and whether it is fast enough?

Lessons in this module

  1. Track quality in production
  2. Manage cost per request and per outcome (this lesson)
  3. Control perceived latency

What you will learn in the course

This lesson is part of the course Ship and monitor an AI feature in production

  • Design the staged launch of an AI feature (feature flag, shadow, canary, percentages) with quantified promotion criteria and a tested kill switch.
  • Define what you log and trace for each model call, what you mask or exclude (personal data), how long you keep it, and with which tool.
  • Set up quality signals in production (user feedback, sampled human review, online evaluations, drift) and connect them to the evaluation set.
  • Manage cost per request and per useful outcome, and perceived latency (budgets, alerts, caching, streaming, response length, model choice).
  • Diagnose AI-specific incidents (drift, provider outage, deprecation, call loop, cost spike) and plan the right fallbacks.
  • Run a model or prompt change with regression evaluations, shadow comparison, staged rollout and a rollback plan.
  • Build the quality, cost and latency dashboard of an AI feature and report to stakeholders with decisions to make.