Evaluate an agent on tasks, trials and trajectories
Lesson 2 of the module "Observe and evaluate an agent" in the course "Design reliable agents and tool calling".
Lesson objective
By the end of this lesson, you will be able to design an agent evaluation: tasks drawn from real cases, graders on the final state, trajectory invariants, multiple trials, pass@k and pass^k metrics, and thresholds set before measuring.
Where it fits
Observe and evaluate an agent
How do I reconstruct what the agent did, and how do I know it is reliable?
Lessons in this module
- Specify a run trace and its metrics
- Evaluate an agent on tasks, trials and trajectories (this lesson)
What you will learn in the course
This lesson is part of the course Design reliable agents and tool calling
- Decide, with explicit criteria, whether a use case calls for a single call, a workflow or an agent, and choose the right pattern.
- Write tool cards the model uses well (name, description, parameters, return, errors, idempotency, risk level).
- Specify an agent's loop (goal, context, stop conditions, budgets, error handling) and judge whether an orchestrator and subagents are warranted.
- Place human approval points and an agent's permissions according to the blast radius of each action.
- Identify prompt-injection paths through tool outputs and specify structural mitigations that do not rely on the prompt alone.
- Specify the trace of a run and the evaluation of an agent (task success, trajectory, cost, latency, consistency over several trials).
Related courses
- Build an AI assistant for your productAdvanced · ~3 hr
- Design a RAG architecture that fits your productAdvanced · ~3 hr 30 min
- Evaluate an AI feature: test sets, metrics and LLM judgesAdvanced · ~3 hr