Agentic AI security controls — the CISO playbook
Published by Zero Hunt, an autonomous AI red team on an on-premise appliance running private AI: automated penetration testing for networks and infrastructure, black-box or gray-box, with a human approving every step that matters.
Short definition
An operational playbook for securing autonomous AI agents and the MCP servers they call — agent identity, tool-use gating, egress monitoring, and a 90-day rollout a CISO can execute.
Why this matters now
In 2025 OWASP added Excessive Agency (LLM06) to its Top 10 for LLM Applications and then shipped a dedicated Top 10 for Agentic Applications; in May 2026 the NSA warned that the Model Context Protocol has proliferated faster than its security model. The stakes are concrete: an over-privileged agent holding standing production credentials and calling tools no one monitors turns a single prompt injection into lateral movement, and the first agent-driven ransomware operations have already compressed that chain to seconds. This playbook is the set of controls that bound what an agent can do before — not after — it becomes the vector.
Key points
- ▸Treat each agent as its own identity: distinct, short-lived, least-privilege credentials — never a shared account with standing production access.
- ▸Tool use is the blast radius: put high-consequence actions (write, delete, pay, deploy) behind an allowlist and human approval, not model discretion.
- ▸MCP servers are a trust boundary, not plumbing: authenticate the server, scope each tool, and log every invocation.
- ▸Prompt injection (LLM01) is the entry; excessive agency (LLM06) makes it an incident — control the second even when you cannot stop the first.
- ▸You cannot govern what you cannot see: inventory every agent — owner, tools, data classes, egress — before the first one reaches production.
- ▸The wire is the backstop: when policy gates are advisory, behavioral egress monitoring catches an agent exfiltrating or beaconing in real time.
Scope — when this playbook fires
Use this playbook when you deploy any AI system that can take actions, not merely generate text: a coding agent that opens pull requests, a SOC copilot that can isolate a host, a customer-facing agent that issues refunds, an RPA flow with an LLM in the loop, or anything wired to tools through the Model Context Protocol or a comparable function-calling layer. The defining property is agency — the system itself chooses which tool to invoke and with what arguments.
In scope: agent identity and credentials, tool-use authorization, MCP server governance, inter-agent trust in multi-agent chains, and the observability needed to prove what an agent actually did.
Out of scope: a pure chatbot with no tool access and no autonomy (that is a data-governance and prompt-injection problem with a lighter control set), and batch inference pipelines that produce text a human then acts on. If the model cannot reach a tool or an API on its own, this playbook is heavier than you need.
The agentic attack surface — what is actually new
Classic application security assumes the code paths are fixed and the inputs are data. An agent breaks both assumptions: it decides at runtime which tool to call, it treats retrieved content and prior conversation as instructions, and in multi-agent designs it trusts the output of other agents as if it were a trusted service.
The threat catalog is now well mapped:
- Excessive agency — OWASP LLM06 in the Top 10 for LLM Applications 2025: the agent holds more permission, autonomy, or tool access than the task requires, so one bad instruction has outsized reach.
- Prompt injection (LLM01) — hostile instructions smuggled through a document, a web page, an email, or another agent's output. You often cannot prevent it; you can bound what it reaches.
- Tool misuse, memory poisoning, privilege compromise, identity spoofing, and cascading multi-agent failures — the threats catalogued by the OWASP Agentic Security Initiative and mapped, for regulator filings, against MITRE ATLAS.
This is not theoretical. The first ransomware operations driven end-to-end by an LLM agent — recon, credential theft, lateral movement, and encryption with no human at the keyboard — compressed a chain that used to take days into one that self-corrects in seconds (see the agentic-AI ransomware containment playbook). Containment designed around human dwell time does not fit. The controls below assume the adversary, or your own compromised agent, moves at machine speed.
Control domain A — identity and least privilege
Goal: every agent is a named, least-privilege principal with no standing access to anything that matters.
- Issue each agent its own identity. Never let an agent run as a human user or share a service account with other agents — you lose attribution and blast-radius control at once.
- Grant just-in-time, time-bound credentials scoped to the specific resource the current task needs. An agent should hold no standing production access between tasks.
- Evaluate authorization at every hop and at the last mile, not only at a perimeter gateway. The MCP server, the API gateway, and any downstream service should each terminate the agent token, check policy per request, and forward only scoped, on-behalf-of credentials.
- Apply the CISA, NSA, and FBI AI data security guidance to the data the agent can read and write: classify it, and let the agent's identity — not the prompt — decide what it may touch.
Checklist for this domain:
- [ ] Every agent has a distinct, inventoried identity.
- [ ] Credentials are short-lived and scoped per task; no standing production access.
- [ ] Policy is enforced per request at every hop, not just the perimeter.
- [ ] Data access is bound to the agent identity and the data classification, not to prompt content.
Control domain B — tool-use gating and MCP governance
Goal: the agent can only invoke tools you have sanctioned, and the actions that can hurt you require a human.
- Maintain an allowlist of tools each agent may call. Default-deny everything else. A new tool is a change that goes through review, not a capability the model discovers at runtime.
- Classify actions by consequence. Read-only retrieval is low-consequence; writing to a database, moving money, deleting data, deploying code, or changing access are high-consequence. Put every high-consequence action behind an explicit human approval step.
- Treat each MCP server as a trust boundary, not internal plumbing. Authenticate the server, pin its identity, give each exposed tool an explicit and minimal contract, and validate both the arguments going in and the content coming back (LLM05, improper output handling applies directly — tool output is untrusted input to the next step).
- Isolate third-party and remote MCP servers from first-party ones. A remote server you do not control is a supply-chain dependency with live reach into your environment; sandbox it, rate-limit it, and log it separately.
Checklist for this domain:
- [ ] Each agent has a reviewed tool allowlist; everything else is denied by default.
- [ ] High-consequence actions require human approval, not model judgment.
- [ ] Every MCP server is authenticated, its tools individually scoped, inputs and outputs validated.
- [ ] Third-party MCP servers are sandboxed and isolated from first-party capability.
Control domain C — observability and egress monitoring
Goal: you can reconstruct what any agent did, and you can see it in time to intervene.
- Log every tool invocation — which agent, which tool, what arguments, what result, under whose approval — to an append-only store the agent cannot edit. This is both your incident timeline and your audit evidence.
- Capture the agent's decision trace where the framework allows it, so a reviewer can see why a tool was called, not just that it was.
- Establish an egress baseline for each agent: which endpoints, which volumes, which cadence are normal. Without a baseline, anomalous beaconing and exfiltration are invisible.
- Watch the network, not only the application logs. Application-layer policy tells you what a well-behaved agent asked for; it does not tell you what a compromised or injected one does once it holds a tool and a route out.
Checklist for this domain:
- [ ] Every tool call is logged immutably with agent, arguments, result, and approver.
- [ ] An egress baseline exists per agent and alerts on deviation.
- [ ] Network-level monitoring covers agent and MCP-server traffic, not just app logs.
- [ ] Audit logs are retained to satisfy the longest applicable regulatory retention.
Evidence checklist and where the controls fail open
Ordered by the gate that consumes it, have ready:
- The agent inventory: every agent, its owner, its tool allowlist, its data classes, its identity.
- The identity register: credential type, scope, lifetime, and issuance path for each agent.
- The tool-and-action policy: the allowlist plus the consequence classification and the approval rules.
- The MCP server registry: each server, first- or third-party, its authentication, its tool contracts.
- The immutable invocation log and the egress baseline with their deviation alerts.
- The red-team results: proof you exercised the agentic surface, with findings and fixes.
The place these controls fail open is the gap between policy and reality. The gates in domains A and B are advisory: they constrain what a well-behaved agent asks for, not what a compromised or prompt-injected one does once it holds a tool and a network path. The backstop is on the wire. An on-premise appliance running AI Traffic Analysis with four parallel inference heads — suspicious traffic, malware classification, attack-type identification, application fingerprinting — at multi-gigabit line rate sees an agent beaconing to a never-before-seen endpoint or exfiltrating through a tool call as behavior, regardless of what the MCP policy permitted, and because it runs locally on the appliance GPU with no external model calls your detection layer is not itself one more agent phoning a cloud LLM. The same autonomous AI red team that validates your other attack paths can exercise the agentic surface directly — chaining a prompt injection into a tool-abuse path, with a human in the loop — so excessive agency becomes a finding you closed in a drill rather than an incident you filed.
Common failure modes
1. The agent runs as a human — or as root. The fastest way to lose control is to hand an agent a privileged user's credentials or a shared admin service account because it was convenient in the prototype. You inherit every permission that account ever accumulated and you lose attribution. Give the agent its own minimal identity before it leaves the prototype.
2. Trusting model-side guardrails as a security control. A system prompt that says do not delete production data is guidance, not enforcement. Prompt injection (OWASP LLM01) exists precisely because instructions and data share a channel. Enforce consequences in the tool layer and the identity layer, where the model cannot talk its way past them.
3. Treating MCP servers as internal plumbing. Skipping authentication and logging on an MCP server because it is on the same cluster is the agentic version of a flat network. The NSA's May 2026 guidance exists because MCP shipped faster than its security model; authenticate and log every server.
4. Human-in-the-loop as a rubber stamp. If every action pops an approval dialog, reviewers approve everything within a day. Scope approvals to the genuinely high-consequence actions so the human attention is spent where it changes the outcome.
5. Pointing your security tooling at a cloud model. If your SOC copilot or your agent monitor itself calls an external LLM API, you have added another channel through which internal data leaves. Keep the detection and analysis layer on infrastructure you control.
6. No egress baseline. Without a model of normal agent traffic, the one metric that reliably distinguishes a working agent from a compromised one — where it talks and how much — is unavailable exactly when you need it.
Governance and regulatory notes
Agentic AI controls do not sit in a regulatory vacuum, and the mapping is what lets you defend the program to a board or an auditor.
- NIST AI RMF. The AI Risk Management Framework and its Generative AI Profile (NIST-AI-600-1) give you the Govern / Map / Measure / Manage structure to organize the controls above and the vocabulary to report them.
- EU AI Act. Where an agent operates inside a high-risk use case, the Act's obligations on risk management, logging, human oversight, and robustness map almost one-to-one onto domains A through C. General-purpose model obligations apply upstream.
- NIS2 and sectoral risk-management duties. NIS2 Article 21 requires appropriate and proportionate technical measures across your systems; an autonomous agent with production reach is squarely within that scope, and supervisors increasingly ask how AI tooling is governed.
- OWASP and MITRE ATLAS are your technical control catalog and your threat-mapping language respectively — cite the specific LLM and agentic risk IDs and the ATLAS techniques in internal risk assessments and in any regulator notification, exactly as you would cite ATT&CK for a conventional incident.
The common thread: the same control evidence — inventory, identity register, tool policy, invocation logs, red-team results — answers every one of these regimes. Build it once, export per framework.
Goes deeper
Want this against your environment?
Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.