Human Approval in AI Agents: Where Should the Checkpoint Go?
Approval works best at the boundary just before a consequential side effect, with an exact operation, clear preview, and server-side verification.
The right approval point is usually after an agent has prepared a concrete action but before that action changes money, customer data, production, or external communication.
Human Approval In Ai Agents Where Should The Checkpoint Go
Human-in-the-loop design is often described as simply asking the user for confirmation. In a production agent, the harder problem is choosing the exact checkpoint and making sure the approved action is the one that actually executes.
{
"operation_id": "email_campaign_42",
"action": "send_email",
"recipients": 214,
"status": "waiting_for_approval"
}Approval should be concrete
Do not approve a vague plan such as “handle this refund.” Approve a specific operation with an operation ID, target record, amount, and expected side effect. The review UI should make the proposed action understandable before execution.
Place the gate before the side effect
Drafting an email does not need the same gate as sending it. Querying account data is different from changing the account. Place approval immediately before the irreversible or expensive operation so earlier safe work remains automatic.
Never turn approval into a blanket token
An approval should expire and be bound to the exact operation. If the agent changes the amount, target, or tool after approval, require a new decision. This prevents stale approvals from becoming general authority.
Make approval usable
Show what matters: affected records, proposed change, reason, confidence signal if useful, and rollback option if available. Avoid forcing a reviewer to inspect raw model messages to understand the action.
Record the decision
Store approver, timestamp, operation ID, policy version, and execution result. A durable approval record lets your audit trail explain who authorized a side effect and whether the system executed the same operation that was reviewed.
Operating table
| Check | Detail |
|---|---|
| Action | Automation |
| Summarize ticket | automatic |
| Draft reply | automatic |
| Send reply | limited |
| Refund | limited |
| Deploy production | limited |
Practical checklist
- Bind approval to operation
- Expire stale approvals
- Show exact side effect
- Verify again on execution
- Audit the decision
Final takeaway
Treat the agent as an untrusted decision-maker inside a controlled software system. The safe architecture is not “trust the model more”; it is to narrow capabilities, enforce policy in code, and make risky actions observable and reversible.
Turning the security principle into an implementation
Security guidance becomes useful when every recommendation maps to a concrete boundary in the application. For Human Approval in AI Agents: Where Should the Checkpoint Go?, begin by listing the assets that could be exposed or changed: customer records, credentials, tokens, production data, private documents, financial actions, and administrative controls. Then identify which component can access each asset and why that access is necessary.
Start with least privilege
Do not give an automated system one credential that can reach the entire application. Create narrow capabilities with explicit scopes. A reporting tool might read aggregated metrics while a billing tool can create an invoice but cannot change account ownership. Separate read operations from write operations and require stronger controls for destructive actions.
Permissions should be enforced by the server or the underlying service, not by instructions inside a prompt. Prompts can explain policy to an agent, but they are not an authorization boundary. Check the authenticated user, resource ownership, role, scope, and requested operation before executing a sensitive tool.
Treat external input as hostile
User messages, uploaded documents, retrieved webpages, emails, tool results, and third-party APIs can all contain instructions that conflict with the application policy. Keep untrusted content distinguishable from system instructions and never let retrieved text silently redefine permissions. If an agent can call tools, validate every tool argument independently.
For database and API access, prefer purpose-built operations over generic capabilities. A function such as get_customer_status is easier to audit than an arbitrary query interface. The narrower the capability, the smaller the blast radius when the model makes a mistake.
Build an incident path
A secure system also needs a response plan. Log authentication failures, permission denials, unusual tool calls, repeated retries, and sensitive operations. Avoid placing secrets or unnecessary personal data in logs. Define how credentials are rotated, how compromised sessions are revoked, and how an unsafe automation can be disabled quickly.
Security review checklist
- Inventory sensitive data and operations.
- Give each service only the permissions it needs.
- Validate authorization on every side-effecting request.
- Keep secrets out of prompts, logs, and client code.
- Treat retrieved and user-provided content as untrusted.
- Add rate limits and abuse controls.
- Make destructive actions reversible where possible.
- Record security-relevant events.
- Test prompt injection and confused-deputy scenarios.
- Define a credential rotation and shutdown procedure.
Security should be designed as a series of enforceable boundaries rather than a final checklist. The strongest implementation is one where an incorrect model response, malicious input, leaked context item, or compromised session still cannot cross the permissions that protect the underlying system.
A founder decision checklist
The final step is to convert the ideas in Human Approval in AI Agents: Where Should the Checkpoint Go? into decisions that can be tested. Start by writing the current state in plain language: what happens today, who owns each step, and where the user or business experiences friction. Then define the desired state and choose one measurement that would show whether the change actually helped.
Before implementation, list the assumptions that could make the plan fail. Separate assumptions about customer behavior from assumptions about technology, cost, timing, and operations. This makes it easier to test the riskiest assumption first instead of spending weeks polishing a solution built on an unverified premise.
During the first release, keep the scope intentionally small. Add logging for the important events, document the expected outcome, and decide what will trigger a rollback. If the workflow involves money, permissions, customer data, or production infrastructure, add an explicit review point before an irreversible action.
After launch, compare the result with the original baseline. Look at a useful cohort rather than only the overall average, record unexpected behavior, and write down the next experiment. A short decision log should capture what changed, why it changed, what happened, and what evidence would justify changing course again.
Use this loop consistently: define the problem, map the workflow, test the riskiest assumption, ship a narrow version, measure the outcome, review failures, and improve the next iteration. That turns a useful idea into a repeatable operating practice instead of a one-time tactic.
Community
What do you think?
0 comments
React to this article
Comments
Trending now
What readers are opening
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?
Continue reading