Caging an autonomous AI agent inside an enterprise tenant
By Matthew Gray on Oct 6, 2026
Autonomous agents that read mail, post in Teams and edit SharePoint are arriving in enterprise tenants faster than the controls around them. An agent is a new kind of non-human user that takes instructions from any text it reads, so it needs the same containment you would give an untrusted contractor with admin tools, designed in before it is switched on.
This is the pattern we design to when a client wants to run a self-hosted autonomous agent inside their Microsoft estate. It comes from scoping and architecture work, not a finished build we can point to, and we say where a control is hard.
What “caging” means
We use “cage” for a short list of properties that must all be true before the agent touches real data:
- it has its own identity, with only the permissions its use case needs
- it runs in an isolated network with no inbound access and allow-listed egress
- a policy layer decides which actions it may take, separate from the model
- every action lands in the SIEM
- consequential actions wait for a human
- someone can stop it, quickly, and has practised doing so
The model is not part of the security boundary. Everything below assumes the model will at some point do something it was told not to.
A dedicated identity per agent
Each agent gets its own identity, never a shared service account and never a person’s credentials. There are two realistic shapes, and the trade-off is worth stating plainly.
- A workload identity (managed identity or federated workload identity) has no password, no MFA question and no licence. It fits agents that call APIs on their own behalf.
- A dedicated user account fits agents that must act as a member of a Teams channel or SharePoint site. It brings real operational work: a licence, an MFA strategy for unattended sign-in, a reviewed Conditional Access exception, and a joiner-mover-leaver process for a user who is not a person. In return, the Microsoft 365 unified audit log records its sign-ins, Graph calls, file reads and posts natively.
Either way, scope permissions to the use case, keep any refresh tokens in Key Vault readable only by the agent’s runtime identity, and alert on that identity signing in from anywhere unexpected.
Where an agent platform has both customers and operators, keep their identity planes separate: customers in one directory, operators from the workforce tenant through cross-tenant access and just-in-time privileged roles. An API gateway validates tokens at the edge, and backends are reached with on-behalf-of or managed identity rather than by passing the user’s token deeper.
Network isolation
The agent runs in a container with no inbound endpoint. Its subnet is default-deny, and egress routes through a firewall with a short list of allowed destinations: identity sign-in, Microsoft Graph, the model endpoint, the vault and the logging ingestion endpoint. Package registries are allowed for the build pipeline and blocked at runtime, so a compromised dependency cannot fetch a second stage. The vault, the log workspace and storage sit behind private endpoints with public access turned off.
The container baseline is ordinary hardening, applied without exceptions:
user: non-root
root_filesystem: read-only # writable paths are named volumes
base_image: minimal, pinned to a release digest
ingress: disabled
egress: firewall only
resources: cpu and memory limits set
sandboxed_execution: true
Resource limits are a security control here: a looping agent should run out of headroom before it floods a channel.
A policy layer the model can’t talk its way past
Many agent runtimes control who may talk to the agent, through per-sender allowlists. Few control what the agent may do. That gap is ours to fill. We put an allow/deny layer between the agent and its tools, with deterministic rules written as Rego or plain JSON, not as prompt instructions. Commands are allow-listed verbs. The agent responds to mentions, not to every message. File drops go to one library and accept only allowed file types. A denied action is logged and can raise an alert.
Actions with consequences, such as sending outside the organisation, changing permissions, or deleting content, pause for human approval in the channel where the request came from. Approval and denial both land in the audit trail.
Prompt injection arrives through tools and data
The attack is rarely someone typing “ignore your instructions” into a chat. It is a document the agent reads, a message it summarises, or a ticket it triages. Every data source the agent can read is an input to its instructions. In threat modelling we cover at least these:
- injection through a chat message, and through file content
- data exfiltration hidden in the agent’s own replies
- a stolen refresh token replayed from outside the network
- the agent calling APIs outside its intended scope
- a malicious package pulled in at runtime
- the system prompt leaking through crafted questions
Each scenario needs a named control: a firewall rule, a policy rule or a specific alert. “We have Defender” is not a mitigation. Per-sender allowlists do not stop injection, because the injected text arrives through a source the agent is allowed to read.
Log everything, and know your latency
We bring five log classes into one workspace that the SOC already uses: container output, identity sign-ins and the unified audit log, network flows from the firewall, application logs, and the agent’s own prompt and action trace. Starting alert rules cover denied egress, unexpected sign-in location, failed sign-in spikes, policy denials, and Graph permission use outside the expected set.
Two cautions. The unified audit log is not real-time; some events take tens of minutes to arrive, so it is for investigation, not for stopping an agent in the act. Near-real-time detection needs the agent’s own trace. Many runtimes don’t export one natively, which can mean building a logging proxy. And model inputs and outputs contain whatever users pasted. Redact personal data before prompts and responses are logged or stored, and turn off request-body logging at the API gateway.
A kill switch you have actually pulled
A kill switch is a runbook, not a button on a slide. It names who can trigger it, which credentials get revoked and rotated, and how the deployment is torn down and rebuilt. We plan at least one full teardown drill before go-live, plus targeted tests of prompt injection, data exfiltration, token replay and sandbox escape. A test passes when a control catches it, not when nothing visibly happens.
The agents you didn’t build
The self-hosted agent you are carefully caging is probably not the only agent in your tenant. Low-code agent builders let business users publish agents with connectors to mail, files and line-of-business systems, often with the author’s own permissions. We treat that estate as an attack surface, using a loop of discover, map, break, control, retest and operationalise. Find every agent, map what it can reach, and try to make it misbehave. Apply the platform’s controls, retest honestly to see whether each closed the gap, dented it or missed it, and make the review recurring.
A note on scoping
A “phase 1” that promises an isolated environment, a running caged agent, monitoring, a kill switch and handover is not a small first step. It is most of a production build. Designing that system and building it are different engagements, and we would rather separate them at the start: a design and a minimal proof first, with the full build scoped from what the design finds. Production-grade infrastructure code, a formal red team and a prompt-trace pipeline are each substantial on their own.
How we can help
If you are planning to run an autonomous agent in your tenant, or suspect you already have more agents than anyone has counted, we can help you design the cage and scope it honestly before anything is switched on. Contact us to talk it through.