Production deployment
Production deployment
Move an agent from prototype to reliable production software with budgets, retries, approvals, and rollback.
flowchart LR
D[Candidate] --> E[Evaluation gate]
E --> S[Small rollout]
S --> O[Observe]
O --> Q{Healthy?}
Q -->|Yes| P[Expand rollout]
Q -->|No| R[Rollback]
P --> O
A[Auth] --> S
B[Budgets] --> S
C[Approval] --> SLesson overview
Production deployment
A demo can succeed once. A production agent must remain bounded, observable and recoverable across thousands of varied runs.
Define the contract
Document supported tasks, unsupported tasks, allowed tools, data boundaries, maximum steps, timeouts, fallback behavior and escalation rules.
Reliability
Use deadlines, bounded retries, idempotency for writes, circuit breakers and durable state for long-running tasks. Never blindly retry an operation that might have already changed state.
Cost controls
Set per-run token, time and monetary budgets. Track cost per successful task. When the budget is exhausted, stop safely or use a deterministic fallback.
Human approval
Require approval for irreversible or high-impact actions such as financial transfers, deletion, external publication and sensitive account changes.
Deployment strategy
Start with a narrow capability set. Test new tools internally, release gradually, use feature flags and maintain rollback paths. A bad prompt or tool version should be disableable without a full emergency rewrite.
Readiness checklist
Before launch verify authentication, tenant isolation, tool schemas, authorization, injection defenses, rate limits, secret handling, trace redaction, evaluation baseline, alerts, incident response and deletion/retention rules.
The goal is not maximum autonomy. The goal is predictable task completion inside explicit boundaries.
Learning path
Theory → Example → Code → Practice → Quiz → Challenge → Completion
Step 1
Theory
Production deployment
A demo can succeed once. A production agent must remain bounded, observable and recoverable across thousands of varied runs.
Define the contract
Document supported tasks, unsupported tasks, allowed tools, data boundaries, maximum steps, timeouts, fallback behavior and escalation rules.
Reliability
Use deadlines, bounded retries, idempotency for writes, circuit breakers and durable state for long-running tasks. Never blindly retry an operation that might have already changed state.
Cost controls
Set per-run token, time and monetary budgets. Track cost per successful task. When the budget is exhausted, stop safely or use a deterministic fallback.
Human approval
Require approval for irreversible or high-impact actions such as financial transfers, deletion, external publication and sensitive account changes.
Deployment strategy
Start with a narrow capability set. Test new tools internally, release gradually, use feature flags and maintain rollback paths. A bad prompt or tool version should be disableable without a full emergency rewrite.
Readiness checklist
Before launch verify authentication, tenant isolation, tool schemas, authorization, injection defenses, rate limits, secret handling, trace redaction, evaluation baseline, alerts, incident response and deletion/retention rules.
The goal is not maximum autonomy. The goal is predictable task completion inside explicit boundaries.
Step 2
Example
Example staged release
Version 1 supports documentation search and ticket drafting. Version 2 adds ticket creation behind approval. Version 3 adds read-only CRM access after evaluation proves tenant isolation. Each capability is introduced separately so failures can be attributed to a specific change.
Step 3
Code
Run budget
const budget = { maxSteps: 8, maxToolCalls: 6, maxCostUsd: 0.05 };
function canContinue(state: State) {
return state.steps < budget.maxSteps &&
state.toolCalls < budget.maxToolCalls &&
state.estimatedCostUsd < budget.maxCostUsd;
}Budget checks are enforced by the runtime, not requested from the model.
Step 4
Practice
Practice
Write a production launch checklist covering identity, data access, tools, retries, idempotency, budgets, approval, evaluation, observability, deployment and rollback.
Step 5
Quiz
1. What should happen when the run budget is exceeded?
2. Why is idempotency important?
Step 6
Challenge
Challenge
Design a staged rollout from internal users to 5%, 25% and 100%. Define the metrics and safety conditions that would stop the rollout at every stage.
Complete every stage
Work through every step in order, then the lesson will be marked complete.
Each chapter and subtopic has its own public URL under /ai-agent.