Module
Module 6 of 7Lesson 2 of 2~19 min

Evaluate an agent on tasks, trials and trajectories

Lesson 2 of the module "Observe and evaluate an agent" in the course "Design reliable agents and tool calling".

Lesson objective

By the end of this lesson, you will be able to design an agent evaluation: tasks drawn from real cases, graders on the final state, trajectory invariants, multiple trials, pass@k and pass^k metrics, and thresholds set before measuring.

Where it fits

Observe and evaluate an agent

How do I reconstruct what the agent did, and how do I know it is reliable?

Lessons in this module

  1. Specify a run trace and its metrics
  2. Evaluate an agent on tasks, trials and trajectories (this lesson)

What you will learn in the course

This lesson is part of the course Design reliable agents and tool calling

  • Decide, with explicit criteria, whether a use case calls for a single call, a workflow or an agent, and choose the right pattern.
  • Write tool cards the model uses well (name, description, parameters, return, errors, idempotency, risk level).
  • Specify an agent's loop (goal, context, stop conditions, budgets, error handling) and judge whether an orchestrator and subagents are warranted.
  • Place human approval points and an agent's permissions according to the blast radius of each action.
  • Identify prompt-injection paths through tool outputs and specify structural mitigations that do not rely on the prompt alone.
  • Specify the trace of a run and the evaluation of an agent (task success, trajectory, cost, latency, consistency over several trials).