Claude Managed Agents vs Self-Hosted Agents
Anthropic's Managed Agents and Claude Agent SDK put the agent loop in different places. The trade-off is operational control versus managed infrastructure.
Anthropic's Managed Agents and Claude Agent SDK put the agent loop in different places. The trade-off is operational control versus managed infrastructure.
Anthropic's current Managed Agents beta separates reusable agents, environments, and sessions, while the Claude Agent SDK runs the agent process in infrastructure you operate.
Anthropic's current Managed Agents beta separates reusable agents, environments, and sessions, while the Claude Agent SDK runs the agent process in infrastructure you operate.
Managed
App → Session → Anthropic environment → tools
Self-hosted
App → Worker → Agent SDK → your sandbox → toolsAnthropic's Managed Agents model separates reusable agent configuration from its execution environment and individual sessions. The agent contains model, instructions, tools, MCP servers, and skills. The environment defines the sandbox and networking. The session performs the work.
A hosted session can handle more of the agent lifecycle, event streaming, and sandbox provisioning. This can be useful for long-running or tool-heavy workflows where operating your own worker process would distract from the product.
With the SDK, your process controls lifecycle, networking, credentials, queues, observability, and the execution environment. This can fit private systems and existing worker infrastructure, but you own failures and scaling.
Even a managed runtime needs a clear boundary around network access and credentials. Keep sensitive business actions as application-owned tools. A sandbox should receive only the credentials and network access needed for the task.
Compare the number of components you must operate, the network controls you need, and how much state you want to own. Do not compare only the number of agent features on a marketing page.
| Metric | Value | Note |
|---|---|---|
| Concern | Managed Agents | Agent SDK |
| Runtime | Provider-managed | You operate it |
| Sessions | Managed | Your application |
| Sandbox | Provider/self-hosted | Your choice |
| Ops work | Lower | Higher |
Build the smallest workflow that proves the customer outcome. Keep tools narrow, business state authoritative, side effects permissioned, and the runtime observable. Agent infrastructure should remove manual work without turning your application into an uncontrolled automation layer.
A sandbox that can reach the whole internet is very different from one restricted to package registries and approved services. Enforce network policy outside the model.
Agent sandbox
+--> package registry allowed
+--> approved MCP allowed
+--> product API allowed
X--> production DB deniedTreat session lifecycle as an operational concern. Decide when a session expires, who can reopen it, what happens after a worker crash, how events are replayed, and how session data is deleted.
Anthropic's current Managed Agents documentation includes scheduled deployments for recurring sessions.
08:00 collect metrics
08:02 analyze anomalies
08:04 save report
08:05 notify ownerA recurring agent should have one defined purpose, a bounded tool surface, and a clear disable path.
The practical question behind this topic is how to turn an AI capability into a dependable product component. Start by defining the job in terms of an input, a useful transformation, and an observable outcome. Avoid designing around the model first. The model is one component inside a workflow that also includes application state, tools, permissions, retries, logging, and user feedback.
For an article about Claude Managed Agents vs Self-Hosted Agents, a useful first exercise is to write the workflow as a sequence of states. Identify what the user provides, what the model needs to know, what information must come from a trusted system, which operations can change data, and what happens when the model is uncertain. This makes hidden assumptions visible before implementation.
A model can interpret a request, classify information, draft a response, or choose between allowed capabilities. It should not become the source of truth for billing, permissions, account ownership, inventory, or destructive actions. Those rules belong in application code. A useful architecture therefore has a clear boundary: the model proposes an action, a typed tool validates it, the application authorizes it, and the system records the result.
This boundary also makes testing easier. Instead of asking whether a model response sounds good, test whether the right tool was selected, whether arguments were valid, whether authorization was enforced, and whether the workflow reached the expected terminal state.
AI systems fail differently from ordinary deterministic services. A response can be syntactically valid but semantically wrong. A tool can time out after the external service has already accepted the request. Retrieval can return stale information. Context can become too large. A retry can accidentally repeat a side effect.
Use explicit limits for turns, latency, token usage, and tool calls. Give side-effecting operations idempotency keys where possible. Store enough trace information to reconstruct the run without storing unnecessary private data. When the system cannot safely continue, return a useful fallback or ask for human intervention rather than silently guessing.
Before launch, collect representative cases rather than relying on a handful of demos. Include normal requests, ambiguous requests, missing data, malformed tool arguments, permission failures, provider errors, and adversarial inputs. Track task success, tool accuracy, latency, cost, and recovery rate. Re-run the same cases after changing the model, prompt, retrieval system, or tool schema.
The goal is not to make the agent appear autonomous. The goal is to make a useful workflow reliable enough that a customer can trust its outcome. Start narrow, keep business rules deterministic, measure the complete job rather than only the model response, and expand the tool surface only when the existing workflow is demonstrably stable.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?