AI Agent Architecture: Models, Tools, Memory, Permissions and Logs
A production agent is a software system around a model. Separate reasoning, tools, state, permissions, approvals, and observability.
A production agent is a software system around a model. Separate reasoning, tools, state, permissions, approvals, and observability.
A useful reference architecture keeps the language model inside deterministic application boundaries. This makes the workflow easier to secure, test, debug, and evolve.
A production AI agent is not a prompt with a longer context window.
It is a software system with several distinct responsibilities.
The model interprets requests and reasons about possibilities.
Tools retrieve data or request actions.
The application enforces authentication and business rules.
State stores information that needs to survive.
Logs explain what actually happened.
Approvals protect high-impact actions.
OpenAI's current agent documentation describes agents as systems that can plan and complete tasks with tools and maintain context across steps. Its platform comparison also makes the location of orchestration and state an explicit architectural decision.
For an indie SaaS, the cleanest mental model is:
request → auth → agent runtime → model → tool → deterministic service → state and logs → response
The model sits inside the boundary. It should not become the boundary.
A language model can be excellent at interpreting ambiguous requests.
It should not become the authority for your business rules.
Suppose a customer asks:
“Cancel my plan and refund the last invoice.”
The model can recognize the intent.
Your server should check the account.
Your billing system should decide refund eligibility.
Your application should enforce who is allowed to execute the change.
The model produces a proposed action. The application turns that proposal into an allowed or rejected operation.
Treat every tool as a small API.
Instead of exposing:
manage_customer
use:
get_customer
list_invoices
create_support_note
request_refund
Each tool should have a clear input schema, authorization rule, error format, and side-effect description.
Small contracts are easier to test and easier to revoke.
OpenAI's current tools documentation describes tools as explicit capabilities attached to model requests. That explicitness is useful because your application can decide exactly which capabilities exist for a given workflow.
Do not create one giant memory bucket.
Use different layers:
Run context for the current task.
Session state for a related interaction.
Durable business data for authoritative application facts.
Long-term memory for selected information that genuinely improves future work.
For example, the customer's current billing plan belongs in the product database.
A recent conversational preference can belong in session state.
A stable formatting preference might belong in memory.
Imagine a research agent that can read documentation and a support agent that can update tickets.
They should not share one broad credential.
Instead:
research agent → read docs, read public product data
support agent → read ticket, read account, write ticket
billing workflow → read invoice, request refund
This gives you a capability map that can be audited.
It also prevents the dangerous pattern where one model credential can reach every internal system.
A human approval step should be represented in the workflow.
For example:
draft email → automatic
send email → approval
create invoice → approval
refund → approval
deploy production → approval
The UI can show the proposed action, affected records, and relevant reason.
The server then verifies the approval before performing the side effect.
A useful agent trace should let you answer:
Who started it?
Which model ran?
Which workflow version ran?
Which tool was selected?
What arguments were accepted?
Was the action authorized?
Was approval required?
Did the tool succeed?
How long did it take?
What was the final outcome?
Avoid logging credentials or unnecessary sensitive content.
OpenAI's current Agents SDK guidance points developers toward traces and evaluation workflows for inspecting and improving runs.
Suppose your pricing system says an account is eligible for a 15% discount.
Do not ask the model to calculate eligibility from raw billing records.
Instead:
model interprets request → pricing service checks rules → service returns decision → model explains
That gives you a testable source of truth.
Tools fail.
Networks fail.
Models produce malformed outputs.
Third-party APIs return partial results.
Users retry requests.
Plan for those cases.
Use validation, bounded retries, timeouts, fallbacks, idempotency keys, and explicit failure states.
A state-changing tool should be safe to retry or should make duplicate execution impossible.
Create a representative test set with normal requests, ambiguous requests, malformed arguments, missing data, approval-required cases, tool failures, conflicting instructions, and adversarial input.
Then compare results across model or prompt changes.
An evaluation set turns “it seems better” into an engineering process.
OpenAI's current cookbook includes examples focused on agent evaluations and improvement loops.
You do not need to build an internal platform first.
A small SaaS can use:
Next.js or Node → authenticated API → agent runner → model provider → narrow tool registry → Postgres or Supabase → structured logs → evaluation dataset → approval UI
The architecture can remain simple until the workload proves that it needs more.
The model should be flexible.
The tools should be narrow.
The business rules should be deterministic.
The permissions should be explicit.
State should be classified.
Logs should be useful.
Risky actions should be approval-aware.
That architecture makes the agent easier to operate because every important responsibility has a clear home.
OpenAI Agents, Agents SDK, tools, and cookbook documentation were checked while updating this article.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?