AI Agent Memory Explained: Context, Sessions and Long-Running Tasks
Memory is not one database. Separate run context, session state, durable facts, and long-term recall.
Memory is not one database. Separate run context, session state, durable facts, and long-term recall.
Reliable agent memory is selective. Keep authoritative facts in the application database, use sessions for active work, and retrieve only the information needed for the current decision.
“Memory” sounds like one feature.
In a production agent, it is usually several different systems.
There is the context the model sees right now.
There is session state connecting related interactions.
There are durable application records.
There may also be a long-term memory layer containing selected preferences or facts.
Mixing them together is one of the easiest ways to create bloated prompts, stale information, privacy problems, and unpredictable behavior.
OpenAI's current agent documentation distinguishes different state approaches across the Agents API, Agents SDK, and Responses API. Its current cookbook also includes separate material on short-term sessions and long-term memory.
What does the model need for this decision?
Maybe the answer is:
current request + three retrieved documents + the last tool result
Once the run ends, much of that information has no reason to stay in active context.
That keeps the working set small.
A session is useful when a user is completing a multi-turn workflow.
Imagine:
Find three competitors.
Then:
Compare their pricing.
Then:
Draft a landing-page section.
The system needs enough continuity for the second and third steps to make sense. It does not necessarily need to carry the entire history forever.
A session is a working area.
The customer's billing plan, account status, product configuration, or subscription limit should have an authoritative source.
Usually that is your application database.
Do not make a generated memory record the source of truth for billing state.
Use:
database → retrieval → relevant fact → model
rather than:
everything we know → giant prompt
This is better for freshness, privacy, cost, and consistency.
A memory is useful when it changes future behavior.
Examples:
Do not save every sentence.
Ask:
Will this information improve a future decision?
If not, it probably belongs in normal conversation history or nowhere.
Imagine 500 support conversations.
Sending all 500 into every request is expensive and noisy.
Instead:
new request → retrieve relevant records → rank results → select the smallest useful context → model
This turns memory into a retrieval problem.
Long-running agents collect messages, tool output, intermediate results, and stale details.
A compact working state can keep:
goal
completed steps
open questions
important facts
recent results
Detailed history can remain in external storage.
OpenAI's current agent materials include guidance around context management and compaction for longer workflows.
The point is not to remember less. It is to remember selectively.
A practical ownership model is:
| State | Owner |
|---|---|
| Current request | request |
| Recent conversation | session |
| Customer/account truth | database |
| Selected preferences | memory |
| Important actions | audit log |
That prevents the same fact from drifting between multiple stores.
A memory table can become a shadow customer database.
Before retaining information, ask:
Would I want this information searchable six months from now?
If not, do not save it by default.
Apply access controls, retention policies, redaction, and deletion.
A generated memory can become stale.
A user may change notification preferences.
A company may change pricing.
A project requirement may be cancelled.
Therefore memory needs a way to be updated, invalidated, or deleted.
When a memory conflicts with authoritative application data, the application record wins.
There are two important failures:
Recall failure: the correct information exists but is not retrieved.
Contamination: irrelevant or stale information is retrieved and influences the answer.
Build evaluation cases for both.
For an indie SaaS:
Postgres or Supabase → authoritative product data
session state → current workflow
memory store → selected durable preferences
retrieval layer → relevant context
agent runtime → reasoning and tools
audit log → important actions
This can be implemented without building a massive memory platform.
Use context for the current task.
Use sessions for active workflows.
Use the database for authoritative facts.
Use long-term memory for selected information that actually helps future work.
Retrieve selectively.
Expire stale information.
Protect sensitive data.
Test both recall and contamination.
An agent does not become smarter by remembering everything. It becomes more useful by retrieving the right information when it matters.
OpenAI's current Agents documentation, cookbook, and state-management materials were checked while updating this article.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?