Place human approvals according to blast radius
Lesson 1 of the module "Human control and security" in the course "Design reliable agents and tool calling".
Lesson objective
By the end of this lesson, you will be able to assess the blast radius of each action an agent takes, place human approval points in the right spots and limit the agent's permissions to those of the user it acts for.
Where it fits
Human control and security
Where must a human approve, with which rights does the agent act, and how do you stop it from obeying booby-trapped content?
Lessons in this module
- Place human approvals according to blast radius (this lesson)
- Counter prompt injection through tool outputs
What you will learn in the course
This lesson is part of the course Design reliable agents and tool calling
- Decide, with explicit criteria, whether a use case calls for a single call, a workflow or an agent, and choose the right pattern.
- Write tool cards the model uses well (name, description, parameters, return, errors, idempotency, risk level).
- Specify an agent's loop (goal, context, stop conditions, budgets, error handling) and judge whether an orchestrator and subagents are warranted.
- Place human approval points and an agent's permissions according to the blast radius of each action.
- Identify prompt-injection paths through tool outputs and specify structural mitigations that do not rely on the prompt alone.
- Specify the trace of a run and the evaluation of an agent (task success, trajectory, cost, latency, consistency over several trials).
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr