Learning/AI Agents — Complete Guide/Lesson 22
Chapter 8·Lesson 1 of 3·10 min

Evaluation

Evaluation

Measure agent quality with datasets, traces, assertions, graders, and regression tests.

Concept diagram
flowchart LR
D[Eval dataset] --> R[Agent run]
R --> T[Trace]
T --> G1[Task grader]
T --> G2[Safety grader]
T --> G3[Tool grader]
T --> G4[Cost + latency]
G1 --> S[Regression report]
G2 --> S
G3 --> S
G4 --> S

Lesson overview

Evaluation

Agent evaluation asks whether the system completed a task correctly, safely and efficiently. A final-answer score is not enough because an agent can produce plausible prose after making an unsafe or incorrect intermediate tool call.

Build a task dataset

Create representative cases from real goals. Include normal requests, ambiguous inputs, edge cases, adversarial content and cases where the correct behavior is to refuse or ask for clarification.

Evaluate layers

Measure task success, factual correctness, tool selection, argument correctness, policy compliance, retrieval quality, step count, latency, cost and escalation rate.

Deterministic graders

Use exact assertions for hard constraints: required tool calls, forbidden tools, schema validity, authorization ordering, maximum steps and approval requirements.

Model-based graders

Use model graders for fuzzy qualities such as helpfulness or explanation quality. Calibrate them against human judgments and monitor grader drift.

Regression testing

Every meaningful production incident should become a permanent evaluation case. Run the dataset whenever prompts, models, tools, retrieval or orchestration changes.

Evaluation is a development loop: establish a baseline, change one variable, compare traces and keep improvements that are supported by evidence.

Learning path

Theory → Example → Code → Practice → Quiz → Challenge → Completion

0/6 done

Step 1

Theory

Evaluation

Agent evaluation asks whether the system completed a task correctly, safely and efficiently. A final-answer score is not enough because an agent can produce plausible prose after making an unsafe or incorrect intermediate tool call.

Build a task dataset

Create representative cases from real goals. Include normal requests, ambiguous inputs, edge cases, adversarial content and cases where the correct behavior is to refuse or ask for clarification.

Evaluate layers

Measure task success, factual correctness, tool selection, argument correctness, policy compliance, retrieval quality, step count, latency, cost and escalation rate.

Deterministic graders

Use exact assertions for hard constraints: required tool calls, forbidden tools, schema validity, authorization ordering, maximum steps and approval requirements.

Model-based graders

Use model graders for fuzzy qualities such as helpfulness or explanation quality. Calibrate them against human judgments and monitor grader drift.

Regression testing

Every meaningful production incident should become a permanent evaluation case. Run the dataset whenever prompts, models, tools, retrieval or orchestration changes.

Evaluation is a development loop: establish a baseline, change one variable, compare traces and keep improvements that are supported by evidence.

Step 2

Example

Example: refund agent

A correct run identifies the order, checks eligibility, requests approval and only then calls the refund API. A test that checks only the final sentence can miss an unsafe refund call that happened earlier, so the trace is part of the expected outcome.

Step 3

Code

Trace assertions

typescript
expect(trace.toolCalls.map(x => x.name)).toEqual(["get_order", "check_refund_policy"]);
expect(trace.approvalRequired).toBe(true);
expect(trace.toolCalls.some(x => x.name === "issue_refund")).toBe(false);

Step 4

Practice

Practice

Create 20 evaluation cases. Label expected outcome, allowed tools, forbidden tools, important facts, maximum steps and whether approval is required.

Step 5

Quiz

1. Why evaluate traces?

2. What should happen after a production failure?

Step 6

Challenge

Challenge

Build a regression report comparing two agent versions on task success, unsafe tool calls, average steps, p95 latency and cost per successful task.

Complete every stage

Work through every step in order, then the lesson will be marked complete.

Each chapter and subtopic has its own public URL under /ai-agent.