Module
Back to courses
AI & ProductExpert

Design reliable agents and tool calling

Decide when an agent is worth it, then design tools, the loop, human approvals, injection defences, traces and evaluation.

53 steps~3 hr 30 minLevel: Expert · Advanced level: for deeper mastery and specialization.

Decide when an agent is worth its cost, and design it to act without surprises: choose between a single call, a workflow and an agent, write tool cards the model uses well, specify the loop and its stop conditions, place human approvals according to blast radius, counter injection through tool outputs, then trace and evaluate runs. You leave with the agent dossier of a real case in your product: tool cards, loop and approval points, ready to review with your engineering team, without writing code.

What you will be able to do

  • Decide, with explicit criteria, whether a use case calls for a single call, a workflow or an agent, and choose the right pattern.
  • Write tool cards the model uses well (name, description, parameters, return, errors, idempotency, risk level).
  • Specify an agent's loop (goal, context, stop conditions, budgets, error handling) and judge whether an orchestrator and subagents are warranted.
  • Place human approval points and an agent's permissions according to the blast radius of each action.
  • Identify prompt-injection paths through tool outputs and specify structural mitigations that do not rely on the prompt alone.
  • Specify the trace of a run and the evaluation of an agent (task success, trajectory, cost, latency, consistency over several trials).

Prerequisites

  • Knowing what an API is and having read a technical spec with an engineering team
  • Having used an LLM for a work task and knowing what a tool call is, at the "I know what it is" level
  • Recommended course: Build an AI assistant for your product (its lesson on tools and actions is taken further here).
  • Recommended course: Get reliable, product-ready AI outputs (output contracts, schemas).
  • This course is not for developers looking for an agent-framework tutorial: tools and loops are described as cards and steps.

Syllabus

What will I design in this course, and in what order?

  1. Objective · By the end of this overview, you will know what you are going to produce (the agent dossier of a real case: decision, tool cards, loop, approvals, mitigations, evaluation) and in what order the six modules get you there.

Does my use case justify an agent, and if not, which simpler pattern should I choose?

  1. Objective · By the end of this lesson, you will be able to classify a use case as a single call, a workflow or an agent using five explicit criteria, and justify that choice to a team.

  2. Objective · By the end of this lesson, you will be able to recognise the five workflow patterns (chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer), match each one to a product need and estimate its cost.

Which tools should the agent get, how should they be described, and what happens when they fail?

  1. Objective · By the end of this lesson, you will be able to design an agent's set of tools (consolidated by task, named unambiguously) and write a complete card for each: description, parameters, return, risk level and application checks.

  2. Objective · By the end of this lesson, you will be able to specify the error messages a tool returns to the agent, make tools with side effects idempotent and describe the expected behaviour on timeout or partial failure.

How does the agent move forward, when does it stop, and does it need several agents?

  1. Objective · By the end of this lesson, you will be able to specify an agent's loop: checkable goal, allowed tools, stop conditions, budgets, failure handling and final output.

  2. Objective · By the end of this lesson, you will be able to judge whether a case warrants an orchestrator and subagents rather than a single agent, and write a subagent's delegation brief.

Where must a human approve, with which rights does the agent act, and how do you stop it from obeying booby-trapped content?

  1. Objective · By the end of this lesson, you will be able to assess the blast radius of each action an agent takes, place human approval points in the right spots and limit the agent's permissions to those of the user it acts for.

  2. Objective · By the end of this lesson, you will be able to spot the paths through which third-party content enters an agent's context via its tools, and specify for each one structural mitigations that do not rely on the prompt alone.

How do I reconstruct what the agent did, and how do I know it is reliable?

  1. Objective · By the end of this lesson, you will be able to specify what a run trace must record to reconstruct what an agent did, and choose the run metrics to track (success, steps, cost, duration, stops, rejected approvals).

  2. Objective · By the end of this lesson, you will be able to design an agent evaluation: tasks drawn from real cases, graders on the final state, trajectory invariants, multiple trials, pass@k and pass^k metrics, and thresholds set before measuring.

How do I assemble the agent dossier of a real case in my product and apply it over the next 30 days?

  1. Objective · By the end of this lesson, you will be able to assemble the agent dossier of a real case in your product (decision, workflow, tool cards, loop, control matrix, injection mitigations, trace and evaluation), self-assess it and plan how to apply it over 7 and 30 days.