How to Build an AI Agent Around the Responses API
Build a focused agent around the Responses API with typed tools, server-side authorization, safe retries, and an evaluation loop.
Build a focused agent around the Responses API with typed tools, server-side authorization, safe retries, and an evaluation loop.
A production Responses API agent is mostly application engineering around a model call: tool contracts, state, authorization, idempotency, observability, and tests.
A production Responses API agent is mostly application engineering around a model call: tool contracts, state, authorization, idempotency, observability, and tests.
const result = await client.responses.create({
model: process.env.OPENAI_MODEL!,
input: message,
tools: [getCustomer, searchDocs]
});Choose a workflow that starts with a clear input and ends with a measurable outcome. Support triage, lead research, invoice review, and document extraction are better starting points than a vague “general agent.”
Give the model the smallest useful context: the user's request, relevant state, and the tools available for this task. Keep credentials and private infrastructure details outside the prompt.
A tool should behave like a normal typed API endpoint. Define the name, purpose, input schema, authorization rule, side-effect level, and output shape. Do not expose a generic execute_sql capability to a business agent.
Tool output should be optimized for the next decision. A customer lookup may only need status, plan, locale, and account ID rather than the entire customer record. Smaller tool results lower context noise and make traces easier to inspect.
Agents fail when the runtime can keep calling tools indefinitely. Set maximum turns, detect repeated failures, and stop when the workflow reaches a terminal state. A final answer should not be the only success condition; record the underlying business outcome too.
| Metric | Value | Note |
|---|---|---|
| Layer | Responsibility | Measure in production |
| Model | Interpretation and reasoning | Measure in production |
| Tool | Typed capability | Measure in production |
| Server | Auth + business rules | Measure in production |
| Database | Source of truth | Measure in production |
| Audit | Record important side effects | Measure in production |
Build the smallest workflow that proves the customer outcome. Keep tools narrow, business state authoritative, side effects permissioned, and the runtime observable. Agent infrastructure should remove manual work without turning your application into an uncontrolled automation layer.
Create a fixture with representative requests, expected tools, approval rules, and expected terminal states.
{
"input": "show invoices for this account",
"expected_tool": "list_invoices",
"requires_approval": false
}| Boundary | Test | Failure to catch |
|---|---|---|
| Auth | wrong user | cross-account access |
| Tool schema | missing field | malformed execution |
| Timeout | slow provider | hanging request |
| Retry | duplicate request | double side effect |
| Approval | forged approval | unauthorized action |
For long workflows, expose status such as “Searching documentation”, “Checking account”, “Draft ready”, and “Waiting for approval”. Users need operational progress, not private reasoning.
A tool should not be allowed to hang the whole request forever.
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 8000);
try {
return await fetch(url, { signal: controller.signal });
} finally {
clearTimeout(timer);
}Use different budgets for different operations. A local database query can have a shorter timeout than an external research call.
The practical question behind this topic is how to turn an AI capability into a dependable product component. Start by defining the job in terms of an input, a useful transformation, and an observable outcome. Avoid designing around the model first. The model is one component inside a workflow that also includes application state, tools, permissions, retries, logging, and user feedback.
For an article about How to Build an AI Agent Around the Responses API, a useful first exercise is to write the workflow as a sequence of states. Identify what the user provides, what the model needs to know, what information must come from a trusted system, which operations can change data, and what happens when the model is uncertain. This makes hidden assumptions visible before implementation.
A model can interpret a request, classify information, draft a response, or choose between allowed capabilities. It should not become the source of truth for billing, permissions, account ownership, inventory, or destructive actions. Those rules belong in application code. A useful architecture therefore has a clear boundary: the model proposes an action, a typed tool validates it, the application authorizes it, and the system records the result.
This boundary also makes testing easier. Instead of asking whether a model response sounds good, test whether the right tool was selected, whether arguments were valid, whether authorization was enforced, and whether the workflow reached the expected terminal state.
AI systems fail differently from ordinary deterministic services. A response can be syntactically valid but semantically wrong. A tool can time out after the external service has already accepted the request. Retrieval can return stale information. Context can become too large. A retry can accidentally repeat a side effect.
Use explicit limits for turns, latency, token usage, and tool calls. Give side-effecting operations idempotency keys where possible. Store enough trace information to reconstruct the run without storing unnecessary private data. When the system cannot safely continue, return a useful fallback or ask for human intervention rather than silently guessing.
Before launch, collect representative cases rather than relying on a handful of demos. Include normal requests, ambiguous requests, missing data, malformed tool arguments, permission failures, provider errors, and adversarial inputs. Track task success, tool accuracy, latency, cost, and recovery rate. Re-run the same cases after changing the model, prompt, retrieval system, or tool schema.
The goal is not to make the agent appear autonomous. The goal is to make a useful workflow reliable enough that a customer can trust its outcome. Start narrow, keep business rules deterministic, measure the complete job rather than only the model response, and expand the tool surface only when the existing workflow is demonstrably stable.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?
Continue reading