Observability
Observability
Trace agent runs end-to-end and connect quality, latency, errors, and cost.
flowchart TD R[Agent run] --> M[Model spans] R --> T[Tool spans] R --> K[Retrieval spans] R --> P[Policy spans] M --> X[Trace store] T --> X K --> X P --> X X --> MET[Metrics] X --> ALERT[Alerts] X --> DEBUG[Debugging]
Lesson overview
Observability
One agent request can produce many model calls, retrieval operations, tool calls, retries and approvals. Ordinary request logs do not explain this causal chain.
Traces
Create one trace for the overall run and spans for model calls, retrieval, policy checks and tool execution. Capture timestamps, status, latency, model/tool identifiers and safe correlation IDs.
Privacy
Prompts and tool payloads can contain secrets or personal data. Redact credentials and unnecessary sensitive fields. Use stable IDs where raw content is not required. Define retention before shipping logs.
Metrics
Track success rate, failure rate, p95 latency, tool error rate, token usage, cost per successful task, steps per run and human escalation rate.
Metrics reveal trends. Traces explain individual failures.
Alerts
Alert on meaningful changes: tool failure spikes, cost anomalies, latency regressions, policy denials or unusual action volume. Do not create alerts for every expected model variation.
Debugging
A useful trace should answer: what did the user ask, what context was supplied, what did the model decide, which tools ran, what did they return, what policy checks happened and why did the run stop?
Learning path
Theory → Example → Code → Practice → Quiz → Challenge → Completion
Step 1
Theory
Observability
One agent request can produce many model calls, retrieval operations, tool calls, retries and approvals. Ordinary request logs do not explain this causal chain.
Traces
Create one trace for the overall run and spans for model calls, retrieval, policy checks and tool execution. Capture timestamps, status, latency, model/tool identifiers and safe correlation IDs.
Privacy
Prompts and tool payloads can contain secrets or personal data. Redact credentials and unnecessary sensitive fields. Use stable IDs where raw content is not required. Define retention before shipping logs.
Metrics
Track success rate, failure rate, p95 latency, tool error rate, token usage, cost per successful task, steps per run and human escalation rate.
Metrics reveal trends. Traces explain individual failures.
Alerts
Alert on meaningful changes: tool failure spikes, cost anomalies, latency regressions, policy denials or unusual action volume. Do not create alerts for every expected model variation.
Debugging
A useful trace should answer: what did the user ask, what context was supplied, what did the model decide, which tools ran, what did they return, what policy checks happened and why did the run stop?
Step 2
Example
Example trace
A single run can show: retrieval 84ms → order lookup 41ms → policy check 3ms → model response 620ms → success. If p95 latency rises, the trace identifies which span changed.
Step 3
Code
Structured event
logger.info("agent.tool.completed", {
runId,
tool: tool.name,
status: "ok",
latencyMs,
requestId
});Do not log raw credentials or unnecessary personal data.
Step 4
Practice
Practice
Define a trace schema with run ID, user/tenant ID, step number, model, tool, status, latency, token usage, cost and policy result. Mark which fields must be redacted.
Step 5
Quiz
1. What does a trace provide?
2. Should raw secrets be logged?
Step 6
Challenge
Challenge
Design an operations dashboard for success rate, p95 latency, cost per successful run, tool failures, approval rate and top failure reasons.
Complete every stage
Work through every step in order, then the lesson will be marked complete.
Each chapter and subtopic has its own public URL under /ai-agent.