agent session management
Agent session management is the ability to watch, quarantine, and end a running AI agent session remotely, so a security or platform team can contain an agent that behaves erratically while the session is still live. It acts on the session itself, rather than on the user account behind it or the tools it calls.
A session is one agent’s continuous run: the context it has accumulated, the credentials it holds, and the connections it has open. Until recently a human sat in front of nearly every session and could close the window when something looked wrong. That assumption no longer holds. Coding agents can work for an hour or more on a single instruction, background agents run in a vendor’s cloud with nobody watching, and a single prompt can produce dozens of tool calls against production systems.
The Cloud Security Alliance’s April 2026 research note on the AI agent governance gap cites a survey of 235 large-enterprise CISOs and CIOs in which 95% doubted they could detect or contain a compromised agent. The same note tells organizations to write agent-specific incident response procedures covering how to revoke an agent’s credentials and how to isolate an agent from the tools it can reach. Agent session management is the set of controls those procedures need in order to be more than a document.
Complete session control requires a component inside the agent harness, the program that runs the agent’s loop. Telemetry and gateways each contribute to session management, but neither can stop a session on its own.
Why do agent sessions need remote management?
An agent session can go wrong in ways its user did not intend and may not notice. Prompt injection is the familiar case: instructions hidden in a web page, a ticket, or a tool result redirect the agent mid-session. Palo Alto Networks’ Unit 42 described a more persistent variant in October 2025, agent session smuggling, in which a malicious remote agent uses an established Agent2Agent (A2A) session to feed covert instructions to a victim agent over several turns. The attack depends on the session carrying state, so the problem lives at the session level.
Agents also misbehave without an attacker. In July 2025, a Replit coding agent deleted a live production database during a declared code freeze, wiping data on more than 1,200 executives and over 1,190 companies despite explicit instructions not to proceed without human approval. The instruction to stop was in the agent’s context the whole time. What was missing was a control outside the agent that could stop it.
The controls most organizations already have act at the wrong granularity for this. Disabling the user’s account or revoking their token stops every session that person has running, along with their email and everything else tied to the identity. Removing a tool from an allowlist stops every agent in the company from using it. Neither lets a responder say “stop this one run, keep the evidence, and leave everyone else alone,” which is usually what a responder needs.
The obvious objection is that credential revocation is enough, and for some organizations it is. If every agent is attended, short-lived, and reaches tools only through credentials the security team can revoke centrally, revoking the credential ends the damage, and the CSA note lists it first for a reason. It falls short in three situations that are becoming the norm:
- Local actions. A coding agent that runs
rm -rfor edits files uses the developer’s own shell and filesystem. No revocable credential is involved. - Token lifetimes. An access token already issued keeps working until it expires unless the downstream system checks revocation on every call, and many don’t.
- Unattended sessions. A background agent with nobody watching keeps running until something stops it, and revoking one credential leaves it free to try the next.
Who is responsible for agent sessions?
Agent sessions usually have three owners, and incidents go badly when nobody has agreed in advance which of them acts.
| Role | Owns | Typical session action |
|---|---|---|
| The person who started the session | The task and its immediate outcome | Stops the agent from the client, approves or rejects prompts |
| Platform or AI enablement team | Harness configuration, hook distribution, the gateway agents use to reach tools | Keeps enforcement points deployed and reporting; applies fleet-wide kill switches |
| Security operations | Detection, triage, and incident response | Quarantines or kills individual sessions; decides on release |
For attended sessions, the user is the fastest responder and usually the right first one. The model breaks down for unattended work. The engineer who kicked off a background agent before lunch is not watching it, and the scheduled agent that runs every night may belong to someone who has since changed teams. Those sessions need a named owner of record, which is one reason an agent inventory comes before session management in most programs.
A workable rule separates pausing from ending. Anyone with an operational stake, including the on-call engineer for the system the agent is touching, can quarantine a session, because quarantine is reversible and preserves evidence. Killing a session and releasing a quarantined one belong to security operations or to the agent’s owner, because both are decisions about the work itself. This mirrors how many organizations already handle production access: broad permission to stop things, narrow permission to restart them.
What is a kill switch for AI agents?
A kill switch is a control that ends an agent session, or a defined class of sessions, immediately and without the agent’s cooperation. An instruction to stop is only a prompt, and the Replit incident shows how little weight a prompt carries once an agent has settled on a course of action.
A complete kill has four effects:
- The agent makes no further model calls in that session.
- The agent’s next tool call is refused, wherever the tool lives.
- Actions the agent takes locally, such as shell commands and file edits, are halted.
- The credentials the session was using are revoked, so a restarted or duplicated session can’t pick up where it left off.
Kill switches differ by scope, and each scope fits a different incident:
- Session. One run is misbehaving. The rest of the person’s work continues.
- Person. A user account is compromised or a departing employee’s agents need to stop. Cutting all of a person’s live connections to tools at the gateway is this scope.
- Agent type. A model update or a bad configuration is causing one kind of agent to misbehave across the fleet.
- Tool. A tool or server is compromised, as in Model Context Protocol (MCP) tool poisoning, and no agent should call it.
A kill is also destructive. The session’s in-progress work is lost, and a kill triggered by a false positive interrupts someone who was doing their job. That cost is the reason most programs put a quarantine step before the kill.
What is session quarantine?
Session quarantine is a control that blocks an agent session’s offending turn and every turn after it until an administrator reviews the session and either releases or ends it. The session stays intact while it waits, so the reviewer can read what the agent was doing and the work can resume where it stopped if the flag was wrong.
Quarantine sits between block and kill on the disposition ladder:
| Disposition | Current action | Later turns in the session | Who resumes |
|---|---|---|---|
| Flag | Proceeds, recorded for review | Proceed | No one needs to |
| Warn | Proceeds after the user acknowledges | Proceed | The user |
| Block | Refused | Proceed | No one needs to |
| Quarantine | Refused | Refused until reviewed | An administrator releases or ends the session |
| Kill | Refused | Session ends | Not resumable |
Two situations make quarantine the right default. The first is unattended work. A warning needs someone to read it, and a background agent has no one, so a blocked action is followed by the agent’s next attempt. Quarantine stops the agent in place until a person says yes or no.
The second is out-of-band analysis. Some of the best signals about a session come from reviewing the whole transcript, which takes longer than the gap between two tool calls. By the time an analysis flags a session, the offending action may already have happened. A block can’t undo that action, but quarantine can stop every action that follows it while the verdict is reviewed.
The trade-off is friction. A quarantine policy tuned too tightly stalls legitimate work, and a queue of quarantined sessions with no one assigned to review them becomes a slower kill. Quarantine also forces a design choice about failure. If the component that checks quarantine status can’t reach the service that holds it, the session either proceeds, which is fail-open and risks missing a quarantine, or stops, which is fail-closed and risks halting work over a network error. Neither answer is free, and the right one depends on what the agent can reach.
How is agent session management implemented?
Agent session management draws on three layers, and each one can do something the others can’t.
| Layer | What it sees | What it can stop | Blind spot |
|---|---|---|---|
| Telemetry (logs, traces, exported transcripts) | What a session did, usually after the fact | Nothing directly | Acts only through another layer |
| Path (an MCP gateway or AI gateway on tool and model calls) | Every call routed through it, with identity attached | The calls it brokers; the credentials it issued | Local actions, unrouted model calls, local tools over stdio |
| Harness (a hook inside the agent loop) | Every prompt and tool call before it runs, including local ones | The agent’s next action and its next turn | Harnesses that don’t expose hooks; hooks a user can remove |
Telemetry tells a responder that a session went wrong. A gateway stops what the session sends through it. The harness stops the session.
Why complete control requires the harness
The agent harness is the program that runs the agent’s loop. It sends the context to the model, receives the action the model chose, and executes it, then repeats. Every action a session takes passes through that loop, including the ones that never touch a network: shell commands, file edits, and local MCP servers that talk to the agent over standard input and output. A component outside the harness can only see and stop the subset of those actions that cross its boundary.
A gateway is the strongest point of control outside the harness, and for tools routed through it the control is real. It can refuse every subsequent call from a session and revoke the credentials it brokered, which is how session kill works at the path layer. But when the gateway refuses a call, the agent’s loop keeps running. The agent can read the refusal, reason about it, and try a different route: a local script, a direct API call with a credential from an environment variable, or a tool the gateway doesn’t front. Refusing calls one at a time is containment of the tools. It doesn’t end the session.
A hook in the harness closes that gap. AI agent hooks are handlers the harness invokes at defined points, such as before a tool runs or when a prompt is submitted, and Claude Code, Cursor, Codex, and VS Code Copilot all expose them. A hook that runs before each action can ask whether this session has been quarantined or killed and deny the action if it has. In Claude Code, a hook can also return a decision that stops the agent from continuing after the current step. That is the one place where a security team can refuse the agent’s next turn, not only its next network call.
The claim cuts both ways. A harness without the other layers is incomplete too. A hook can’t revoke a token that’s already been issued, and a hook on a laptop runs on a machine the user controls. The architecture that gives complete control uses all three layers:
- Telemetry decides. Session transcripts and call logs feed the detection that flags a session.
- The path revokes. The gateway refuses the session’s calls and invalidates its credentials server-side, where the user can’t interfere.
- The harness halts. The hook refuses the next action, local or remote, and stops the loop.
For those layers to act on the same session, they need a shared session identifier. If the gateway knows the session by one ID and the hook knows it by another, a quarantine applied in one place won’t be enforced in the other.
Where harness participation isn’t available
Not every agent exposes a harness a customer can hook. Hosted assistants and vendor-run background agents run their loop on infrastructure the customer doesn’t control, and the session controls available are whatever the vendor exposes, if any. Dropping a live session or clearing an agent’s persistent memory, for example, is possible only where the platform provides a way to do it. For those agents, the realistic target is path control plus telemetry, and the honest description of the result is containment of tools, not control of the session.
Hooks on managed endpoints have their own caveats. Several harnesses treat a failing hook as non-blocking, which keeps a broken script from stopping a developer’s work but also means a kill that depends on an unreachable backend may not fire unless the hook is written to fail closed. A hook that lives in a user’s own configuration file can be edited out by that user, so hooks meant for enforcement have to be distributed as managed configuration, for example through a harness’s enterprise settings or an MDM profile.
Where to start
Ranked roughly by effort, from least to most:
- Make sessions visible. Get a live list of agent sessions with an owner and a session identifier for each. Without it, there is nothing to quarantine.
- Put tool calls on a path. Route agent tool calls through a gateway bound to directory identity, so a session’s access can be refused and its credentials revoked centrally. This overlaps with agent access management, which decides what a session may call in the first place.
- Add the harness. Distribute hooks as managed configuration to every harness that supports them, and have them check session status before each action.
- Write the ladder and the owners. Decide which signals map to flag, warn, block, quarantine, and kill, and who can apply and release each one.
- Pull the kill switch on purpose. Run a kill against a test session and confirm that all four effects happen. A kill switch nobody has tested is an assumption.
A note on Speakeasy
Speakeasy builds an AI Control Plane that includes both an MCP gateway and agent hooks, the path and harness layers described above. If you’re working out who in your organization should hold the kill switch for agent sessions, we’d be interested to compare notes. Talk to us.