There is a point in the adoption of any big technology shift where the security implications stop being theoretical and start being urgent. Agentic AI is at that point now, and most organisations are behind the threat curve.

The shift itself is easy to describe. For the past two years, enterprise AI has mostly answered questions: draft this email, summarise this document, explain this code. A human stayed in control, and the AI advised. The next wave, already live in thousands of organisations, is different. Agentic systems act. They browse the web, call APIs, read and write files, send messages, execute code and hand work to other agents, autonomously, at machine speed and often with nobody reviewing each step. That changes the security calculus.

Why traditional security fails here

The core problem is intent opacity. A network packet generated by an AI agent is structurally identical to one generated by a legitimate application. Standard protocols, real endpoints, valid payloads. Your firewall cannot tell the difference, and neither can your SIEM, your DLP tooling or your WAF. The difference between an agent doing its authorised job and one an attacker has hijacked lives entirely in the semantics of what it is doing, and traditional security tools do not understand semantics.

And the permissions make it worse. A customer service agent might read your CRM and initiate refunds. A code assistant might write to your repo and trigger CI/CD. Each grant is reasonable on its own. Put them together and a compromised or manipulated agent is an insider threat running at a speed and scale no human attacker could match.

The threat you need to understand first

One risk needs covering before any discussion of controls, because it is the most severe and the least solved: prompt injection.

Prompt injection is the AI equivalent of SQL injection. An attacker embeds malicious instructions in content the agent will process: a web page it browses, a document it reads, an email it summarises, and those instructions hijack the agent's behaviour. The agent believes it is following its legitimate instructions. It is following the attacker's. And unlike SQL injection, there is currently no deterministic fix. No patch removes it. Defence has to come in layers.

Which is why agentic AI security cannot be a checkbox exercise. It needs defence in depth, continuous monitoring, and an acceptance that some of this risk can only be managed, never removed.

Governance is the foundation

Every technical control below depends on governance decisions most organisations have not made yet. Someone has to decide what agents are allowed to do, what data they can touch, which actions need a human, and what happens when it goes wrong. Until then, technical controls have no policy to enforce.

In practice that means four things. Establish agent identity and authorisation scope, treating agents the way you would treat privileged service accounts, with explicit documented permissions and regular review. Define human in the loop thresholds for actions too consequential to run autonomously. Mandate immutable audit logging of agent decisions and tool calls, because regulators and legal will eventually want to know exactly what an AI did and why. And build a kill switch: a defined way to stop an agent immediately, plus a process for rolling back whatever it changed.

World Wide Technology's AI Readiness Model for Operational Resilience (ARMOR) gives organisations a structured, practical way to build exactly this. Its AI Governance, Risk and Compliance domain, one of seven covering the full breadth of AI security, treats governance as a set of capabilities embedded throughout the AI lifecycle rather than a compliance exercise, with risk management, governance frameworks and compliance alignment built in from the start instead of retrofitted after deployment. ARMOR also maps to the leading regulatory and standards frameworks, which matters if you need agentic AI governance that satisfies security requirements and regulators at the same time. For those organisations the domain is a practical starting point connecting strategic intent to operational execution.

The technical controls that actually matter

Everything in this section rests on one premise: assume the workload is hostile. Not because your agents are badly designed, but because any of them, through prompt injection, a compromised dependency, a misconfigured permission or a supply chain attack, could be acting against your interests at any moment. Controls built around the question of what a compromised workload can actually do hold up far better than controls built around preventing compromise, because the first question accepts a reality the second ignores.

Give every agent its own identity. Shared API keys are one of the most common mistakes in agentic deployments, and one of the most dangerous. Every agent instance needs its own credential: unique, cryptographic, short lived, bound to its deployment context, ideally on a workload identity standard like SPIFFE. If a credential is compromised, you revoke that agent and only that agent, and your audit trail shows exactly what the identity did. Under an assume-hostile architecture, credential isolation is a containment mechanism, not a convenience. When an agent turns adversarial, you need to cut it off immediately without collateral damage across the estate. ARMOR's Infrastructure Security domain makes identity its first principle, moving organisations from manual access controls to intelligent, risk-based identity management with zero standing privilege, and carries detailed implementation guidance for AI environments.

Then harden what the agent runs on. Container images, base operating systems, runtime libraries and AI frameworks all carry vulnerability backlogs attackers understand well. Chasing CVSS scores produces remediation queues that look busy and reflect nothing. What matters is exploitability in context. A critical vulnerability in a library no code path reaches is a lower priority than a medium severity flaw in a component your agent hits on every external request. Agent runtimes are trust boundary workloads: they consume external content, call external services and act on what comes back, so they deserve the fastest remediation SLAs in your estate, days not quarters, with automated scanning in the CI/CD pipeline and policy as code that rejects failing images before production. This is ARMOR's vulnerability management principle in the Infrastructure Security domain: risk-based vulnerability management and rapid remediation.

Configuration is the weaker discipline in most estates, and it is the complement to patching. A fully patched agent with permissive settings is often trivially exploitable. Run agents without root, with read-only filesystems where possible. Drop capabilities to the explicit minimum. Apply seccomp and AppArmor profiles that constrain which system calls the process can make. At the orchestration layer, disable automatic mounting of service account tokens, enforce pod security standards and use network policy to bound the blast radius of a compromised pod before any application layer control gets involved. The logic is simple. A hostile workload that cannot write to the filesystem cannot persist. One that cannot make arbitrary outbound connections has to beat your egress controls to get anything out. Configuration hardening defines what a compromised agent can actually accomplish. ARMOR covers this under the secure build principle in Infrastructure Security, hardened build standards and automated compliance checks.

DNS filtering earns its place as the baseline control. Agents that browse, call APIs and talk to external services generate network behaviour traditional controls were never designed for. DNS filtering blocks connections to malicious, unauthorised or unexpected domains before any harmful communication happens, and because it sits beneath the agent's own logic it cannot be reasoned away from inside the agent's context, which makes it resilient to prompt injection. The query logs give you an audit trail of agent behaviour into the bargain. Even a fully hijacked agent cannot exfiltrate to an attacker-controlled domain your DNS layer refuses to resolve. Simple to deploy, scales well, do it early.

Put an AI security gateway in front of model traffic. All agent traffic to external model providers and tool APIs should flow through an inline proxy that enforces authentication, inspects content semantically, rate limits and logs in detail. Think of it as a next-generation firewall that understands what AI payloads mean, not just how they are structured. The gateway is where a compromised agent's attempts to exfiltrate through model API calls, invoke unauthorised tools or reach external infrastructure become visible and stoppable, even after its own guardrails have been bypassed. ARMOR's Model Protection domain has model gateways sitting alongside access controls and encryption as the primary mechanisms for protecting models in use.

Apply guardrails at every boundary, and be clear-eyed about what they are. Unlike a firewall rule or an ACL, AI guardrails are probabilistic. They will not catch every threat every time. A well-built prompt injection may get through. A leakage attempt may be missed. Better models will not fix this; it is a property of how these systems work, and it is exactly why defence in depth matters so much here. No single guardrail, or any other single control for that matter, should be your last line of defence, because none of them can be.

Input guardrails inspect what enters the agent's context, catching injection attempts and inputs designed to push the agent past its authorised scope. Output guardrails inspect what the agent produces before it reaches downstream systems, catching leakage and responses carrying instructions aimed at other agents. Tool call guardrails are the layer that matters most: they validate that every action an agent attempts is consistent with its declared task and authorised scope, not merely technically permitted. Stack all three knowing each will sometimes fail, and knowing a hostile workload will probe for exactly those failures.

The NeuralExec attack against Apple Intelligence in April 2026 shows the failure mode in the wild. Attackers embedded invisible Unicode characters inside ordinary documents and used them to manipulate the model's internal reasoning in ways its safety guardrails could not detect. Probabilistic controls plus excessive agency is a dangerous combination, and network layer enforcement, least privilege permissions and layered detection are not optional extras on top of a guardrail first strategy.

ARMOR's Model Protection domain frames guardrails as a runtime enforcement layer designed in from the start rather than a single control, with red teaming and adversarial testing feeding back into guardrail policy and retraining.

Enforce zero trust at the network layer. Agent compute belongs in dedicated segments with minimal, explicitly defined egress. An agent permitted to call two external services should have firewall policy permitting exactly those two destinations. This is not a substitute for application layer security, it is the backstop when application layer controls fail, and it is one of the most reliable backstops available because nothing the agent processes can influence it. A prompt-injected agent told to send data to an attacker-controlled domain fails at the network layer if your egress policy is sound, whatever happens above. ARMOR's Infrastructure Security domain builds zero trust into its network security guidance, pairing segmentation and boundary protection with continuous network visibility for real-time threat monitoring, and provides the maturity benchmarks and implementation steps to get there.

And baseline behaviour. Every agent type has a characteristic pattern: which tools it calls, in what order, how often, against what data. Log it and alert on deviation. An agent that suddenly reads HR data when its normal scope is customer records is a high-priority signal whether or not it technically has permission. Assuming hostility makes behavioural monitoring more useful, not less; anomalies stop being odd edge cases for the backlog and become confirmation that something is already wrong. ARMOR ties this to the lifecycle traceability principle in Model Protection: keep audit trails, track the lineage of AI artifacts and make sure every action an agent takes can be attributed, reviewed and, where needed, challenged.

If the worst does happen, ARMOR's Cyber Resilience domain covers the response side, integrating people, process and technology to protect business operations and critical services, and treating resilience as a core business discipline rather than a technology bolt-on. Security architecture, data protection systems, integrated governance, incident response and recovery planning all sit in that domain.

The platforms are not all equal

Custom agents built on OpenAI's Custom GPTs, Google's Gemini Gems, Anthropic's Claude Projects or Perplexity Spaces are not interchangeable from a security perspective. A Custom GPT with external Actions or a Gem with Workspace extensions presents a substantially larger attack surface than a Claude Project holding document context with no tool calling. Map controls to what each platform's agents can actually do, particularly whether they can call external APIs and reach organisational data stores, and apply audit logging, admin governance of what can be created and shared and data access scope restriction consistently across every platform in the estate.

Where to begin

The organisations that come through this well will treat agentic AI security as an architectural concern from day one rather than a retrofit, and that means building for compromise, not just against it. Assume some of your agents will be hijacked through prompt injection, a vulnerable dependency or a misconfigured permission, and let that assumption set the priorities. Inventory every agent and every tool it can call, then review the map for what a hostile version of each agent could accomplish, not just what it is authorised to do. Kill shared credentials immediately; one compromised agent should not be able to impersonate every other agent you run. Get a gateway proxy in front of external model API traffic. Cut blast radius with segmentation and DNS filtering, the controls that sit beneath the agent's logic and cannot be talked out of the way by an attacker already inside. Set human-in-the-loop thresholds before you need them under pressure, and set each one by asking how much damage the action would enable if an adversary controlled the agent right now. And stand up a governance body with the authority and cross-functional representation to make binding decisions about what your agents may do.

WWT can help at every stage. ARMOR provides the domain-by-domain guidance and maturity benchmarks to work out where your agentic AI security posture stands and where it needs to be. WWT's AI security assessment services turn that framework into a concrete view of your specific environment, finding the gaps in governance, identity, network controls and guardrail coverage before an attacker does. On gateway selection, WWT brings vendor-agnostic coverage of the AI security gateway market to match the enforcement architecture to your workload profile and risk appetite. The AI Proving Ground in WWT's Advanced Technology Center gives teams hands-on lab capability to test and validate agentic AI security controls before anything reaches production. And on governance design, agent authorisation scope, human-in-the-loop thresholds and the policy frameworks that give technical controls something to enforce, WWT's consulting practice works across technical, legal and business domains to build programmes that are practical, auditable and aligned with whatever regulatory environment you operate in.

None of the principles are new. Zero trust, least privilege, defence in depth and assume breach have underpinned good practice for years. What is new is the domain, the speed it is expanding at, and the cost of applying those principles late. Agentic AI is already in production. The only real question is whether your architecture assumes the worst before the worst happens, or after.