Evaluation
Evaluation
Measure agent quality with datasets, traces, assertions, graders, and regression tests.
flowchart LR D[Eval dataset] --> R[Agent run] R --> T[Trace] T --> G1[Task grader] T --> G2[Safety grader] T --> G3[Tool grader] T --> G4[Cost + latency] G1 --> S[Regression report] G2 --> S G3 --> S G4 --> S
Lesson overview
Evaluation
Agent evaluation asks whether the system completed a task correctly, safely and efficiently. A final-answer score is not enough because an agent can produce plausible prose after making an unsafe or incorrect intermediate tool call.
Build a task dataset
Create representative cases from real goals. Include normal requests, ambiguous inputs, edge cases, adversarial content and cases where the correct behavior is to refuse or ask for clarification.
Evaluate layers
Measure task success, factual correctness, tool selection, argument correctness, policy compliance, retrieval quality, step count, latency, cost and escalation rate.
Deterministic graders
Use exact assertions for hard constraints: required tool calls, forbidden tools, schema validity, authorization ordering, maximum steps and approval requirements.
Model-based graders
Use model graders for fuzzy qualities such as helpfulness or explanation quality. Calibrate them against human judgments and monitor grader drift.
Regression testing
Every meaningful production incident should become a permanent evaluation case. Run the dataset whenever prompts, models, tools, retrieval or orchestration changes.
Evaluation is a development loop: establish a baseline, change one variable, compare traces and keep improvements that are supported by evidence.
Learning path
Theory → Example → Code → Practice → Quiz → Challenge → Completion
Step 1
Theory
Evaluation
Agent evaluation asks whether the system completed a task correctly, safely and efficiently. A final-answer score is not enough because an agent can produce plausible prose after making an unsafe or incorrect intermediate tool call.
Build a task dataset
Create representative cases from real goals. Include normal requests, ambiguous inputs, edge cases, adversarial content and cases where the correct behavior is to refuse or ask for clarification.
Evaluate layers
Measure task success, factual correctness, tool selection, argument correctness, policy compliance, retrieval quality, step count, latency, cost and escalation rate.
Deterministic graders
Use exact assertions for hard constraints: required tool calls, forbidden tools, schema validity, authorization ordering, maximum steps and approval requirements.
Model-based graders
Use model graders for fuzzy qualities such as helpfulness or explanation quality. Calibrate them against human judgments and monitor grader drift.
Regression testing
Every meaningful production incident should become a permanent evaluation case. Run the dataset whenever prompts, models, tools, retrieval or orchestration changes.
Evaluation is a development loop: establish a baseline, change one variable, compare traces and keep improvements that are supported by evidence.
Step 2
Example
Example: refund agent
A correct run identifies the order, checks eligibility, requests approval and only then calls the refund API. A test that checks only the final sentence can miss an unsafe refund call that happened earlier, so the trace is part of the expected outcome.
Step 3
Code
Trace assertions
expect(trace.toolCalls.map(x => x.name)).toEqual(["get_order", "check_refund_policy"]);
expect(trace.approvalRequired).toBe(true);
expect(trace.toolCalls.some(x => x.name === "issue_refund")).toBe(false);Step 4
Practice
Practice
Create 20 evaluation cases. Label expected outcome, allowed tools, forbidden tools, important facts, maximum steps and whether approval is required.
Step 5
Quiz
1. Why evaluate traces?
2. What should happen after a production failure?
Step 6
Challenge
Challenge
Build a regression report comparing two agent versions on task success, unsafe tool calls, average steps, p95 latency and cost per successful task.
Complete every stage
Work through every step in order, then the lesson will be marked complete.
Each chapter and subtopic has its own public URL under /ai-agent.