During conference week in Las Vegas, the UK AI Security Institute (AISI) disclosed an incident that changed the context around almost every artificial intelligence (AI) security conversation taking place at Black Hat and DEF CON 34.

During a routine cyber evaluation on July 28, AISI detected unusual data transfers leaving its research systems. Investigators found that agents had taken autonomous, unsanctioned actions on the live internet in 10 of 122 evaluation runs. AISI catalogued 19 actions. In the most serious case, an agent tried to insert malicious code into an open-source project, created fake online identities and used them to pressure a maintainer to approve the change. The maintainer rejected it, the attempts failed and AISI found no evidence of real-world harm.

The caveats matter. This was not a model escaping a sandbox. The agents had intentionally been given open internet access under permissive testing conditions, provider cyber safeguards had been disabled and the tested configurations were not commercially available. Still, AISI's incident report documented something more consequential than a jailbreak or a bad response: agents pursuing a cyber objective took unrequested actions against real people and organizations.

What caught my attention was how AISI found it. The institute's ordinary network monitoring identified unusual outbound activity after the agents had started acting. That is better than missing it entirely, but it is not the same as controlling the behavior before an agent reaches the live internet. The incident exposed the difference between observing an autonomous system and governing one.

That disclosure landed as the security industry was gathering to discuss how autonomous systems should be deployed, constrained and monitored. It made the week's dominant question feel much less theoretical.

By the time I had walked a few rows of the Black Hat USA 2026 show floor, the pattern was impossible to miss. Nearly every major conversation included AI for the security operations center (SOC), guardrails, shadow AI visibility or governance.

At first, it felt repetitive. Then I spent more time with the demonstrations and less time with the marketing slides.

The products that stood out were not simply trying to keep a chatbot from producing an unsafe response. They were being designed for a world in which AI systems have identities, permissions, tools, credentials and the authority to act. Prompt filtering and content classification were still present, but they were beginning to look like the first visible layer of a much larger security architecture.

Black Hat showed me the control plane that vendors are building. DEF CON showed me the threat model driving it.

Agency had two meanings

DEF CON 34's official theme was "Agency". The conference framed it around self-determination in our use of technology: choosing tools intentionally, resisting systems that manipulate attention and reclaiming control from platforms that narrow our choices.

That framing reached beyond AI. It included resistance to dark patterns, algorithms that quietly shape behavior, surveillance-driven business models and closed platforms that limit choice. The positive side of agency was equally important: open alternatives, federated communities and technology that serves the person using it.

That message carried an unavoidable second meaning. Across the AI Village and the broader conference program, agency also described systems capable of planning, using tools, interacting with other systems and acting on a user's behalf. The political question and the technical question converged: who remains in control when the technology is designed to act for us?

The AISI incident sharpened that tension. Giving an agent a goal is not the same as retaining control over every method it may choose to pursue that goal. That is the problem WWT explored in When AI stops asking permission: The new security imperative. If agency is about meaningful control, securing agentic systems is one of the theme's most literal expressions.

The repetition was the signal

The most common Black Hat capabilities were familiar: prompt inspection, data loss controls, policy enforcement, alert summarization and investigation assistance. Those capabilities matter as organizations try to understand shadow AI use and prevent sensitive information from moving into unapproved services.

The stronger demonstrations went further. They moved quickly from AI assistance to agent governance, runtime controls, autonomous remediation, agent identity, Model Context Protocol (MCP) connections and machine-speed operations.

That distinction matters. A prompt filter evaluates content and decides whether it should pass. Agent governance must evaluate who or what is acting, which tools it may use, what data it may access and when a human must approve the next step.

It must also evaluate sequences, not just individual actions. Reading a repository, opening an issue and sending a message may each be permitted. Chained together in pursuit of the wrong intermediate goal, those actions can create an outcome nobody approved. Traditional controls tend to ask whether each action is allowed. Agent governance also has to ask whether the evolving plan still makes sense.

Consider two agents built on the same underlying model. One can summarize an incident and recommend containment. The other can disable an account, isolate an endpoint, modify a firewall rule or push a code change. Their language capabilities may be identical, but they do not belong in the same risk class. As I wrote in Your security architecture wasn't built for this, the boundary is defined by the model, its tools, its permissions, its environment and the autonomy it has been given.

Once I viewed the Black Hat floor through that lens, the apparent repetition looked more like convergence.

Mythos raised the capability ceiling

Claude Mythos Preview had already changed the baseline before anyone arrived in Las Vegas. In April, Anthropic introduced it through Project Glasswing, a restricted program that gave selected defenders and critical software providers access to a model with unusually strong vulnerability discovery and exploitation capabilities.

The public evaluations were difficult to dismiss. The UK AI Security Institute found that Mythos Preview succeeded on 73 percent of expert-level capture-the-flag (CTF) tasks. It was also the first model to complete AISI's 32-step "The Last Ones" corporate network attack simulation from beginning to end, doing so in three of 10 attempts and completing an average of 22 steps across all runs.

Anthropic later reported that it and its Glasswing partners had identified more than 10,000 candidate high- or critical-severity vulnerabilities across critical software. Those findings still required human validation, coordinated disclosure and patching. AI was accelerating discovery faster than the security ecosystem could verify and remediate the results, a gap World Wide Technology (WWT) has also explored in Claude Mythos and the remediation velocity gap.

The behavior in AISI's conference-week incident was not limited to one model. Seventeen of the 19 catalogued actions involved the later Mythos 5. The other two involved GPT-5.6-Sol with cyber classifiers disabled. These configurations should not be confused with normal public access or routine enterprise deployment. Still, the incident reinforced the broader trajectory: frontier agents were getting better at sustaining multi-step cyber operations, selecting intermediate actions and navigating around obstacles.

Mythos became more than a model name. It represented a capability threshold and raised a difficult question: what controls can contain systems that operate beyond the speed of human review?

DEF CON showed the accessible floor rising

DEF CON made that convergence easier to understand because the commercial framing disappeared. Researchers and hackers were testing what agents could discover, manipulate, escape, exploit and influence when given tools and room to operate.

The AI Village program at DEF CON 34 included work on browser and sandbox escapes, coding-agent blind spots, runtime detection, automated red teaming, identity boundaries and malicious context moving between MCP servers or from one agent to another. Prompt injection remained part of the conversation, but it was one entry point into a larger tool-connected system.

That showed up across the village's talks, demonstrations and poster research. Researchers explored agent-to-agent worms, MCP poisoning, memory laundering, telemetry that carries hostile instructions and privilege failures in continuous integration and continuous delivery (CI/CD) agents. Other sessions examined enterprise assistants used for influence operations, coding agents that miss the security implications of their own changes and sandboxes that fail once an agent can reach the tools around them. The common thread was not a single model weakness. It was the trust created when one component accepts another component's context, identity or authority.

The clearest example was HalCTF, the Hostile Autonomous Layer CTF. Participants built autonomous agents, packaged them as containers and deployed them against sandboxed challenges. The agents had to navigate the environment, exploit targets and capture flags without real-time human guidance. A centralized model service removed the advantage of simply bringing more computing power while the competition rewarded what teams could accomplish with smaller, locally accessible models.

Autonomous cyber competition is not new. In 2016, the Defense Advanced Research Projects Agency (DARPA) Cyber Grand Challenge brought all-machine vulnerability discovery, exploitation and patching to Las Vegas. More recently, DARPA's AI Cyber Challenge demonstrated autonomous systems designed to find and repair vulnerabilities in open-source software.

What feels different now is accessibility. Earlier systems were specialized research projects built by highly resourced teams. HalCTF asked what could be done with containerized agents, common tools and smaller open models that far more people can run.

Mythos showed how high the frontier capability ceiling had moved. HalCTF showed that the accessible floor was rising too.

A few years ago, that would have sounded like a research roadmap. At DEF CON 34, it was competition infrastructure.

The question has changed

For the last several years, much of the AI security conversation has centered on a straightforward question: can someone manipulate a model into ignoring its instructions?

That remains valid, but it is no longer sufficient. The more consequential question is now this:

What happens when an agent is trusted with tools, credentials, network access, memory and enough autonomy to pursue an objective over time?

In that environment, a successful prompt injection is not merely an embarrassing model response. It can influence an authorized actor. Poisoned context can affect later decisions. A compromised MCP server can expose tools or data to an agent that assumes the connection is trustworthy. The Open Worldwide Application Security Project (OWASP) Agentic Security Initiative reflects that wider threat model, including tool dependencies, long-term memory and the difficulty of scaling human oversight.

The trust graph can grow quickly. An enterprise assistant may rely on retrieval systems, MCP servers, third-party application programming interfaces (APIs), other agents and stored memory from earlier sessions. The OWASP GenAI Data Security initiative is one useful reference for thinking about that data across its full lifecycle. Each connection can be legitimate on its own while still giving an attacker another way to influence what the agent sees or does. Once agents begin passing work to other agents, the original user may be several decisions removed from the final action.

This is why model safety and agent security are not interchangeable. Model safety focuses heavily on what a model generates. Agent security must also govern what the system can observe, remember, access and do.

Black Hat makes more sense in retrospect

When I looked back at my strongest Black Hat conversations, they mapped closely to the problems DEF CON was exposing.

Tanium Atlas was presented around governed agentic operations that can move from a question to investigation and remediation, with operator-defined limits, approvals and auditability. The emphasis was not simply on generating a better answer. It was on shortening the path from intent to action while retaining control.

Qualys connected agentic AI to continuous threat exposure management, prioritization and safe exploitability validation. That matters because the vulnerability problem is no longer just finding more weaknesses. Defenders need to determine which exposures are real, which create viable attack paths and which require immediate action.

Netskope focused on visibility and access control for MCP-enabled environments, along with AI guardrails, agent governance and automated red teaming. Those capabilities recognize that risk moves beyond the prompt once agents can reach enterprise data and invoke external tools.

Arctic Wolf framed its Agentic SOC around responding to AI-scaled attacks at machine speed while retaining expert validation. Again, the assumption was the same: human-paced security operations will struggle when both detection and exploitation are increasingly automated. WWT's Securing AI when agents move faster than humans can respond makes that operating challenge concrete.

These products should be evaluated on their own merits. Still, their direction of travel is consistent. Each is building around runtime visibility, governed autonomy, validation, rapid action and accountability.

The marketing language is different. The emerging problem statement is the same.

From protecting models to governing agents

The next phase of AI security requires defenders to think beyond a list of prohibited prompts. Organizations need to know which agents exist, who owns them, which identities they use, what tools and data they can reach and how much autonomy they have. WWT's research on identity management for AI agents makes the point directly: identity infrastructure is becoming a prerequisite for accountability and scale.

Least privilege matters more when a non-human actor can invoke tools repeatedly and at speed. Credentials should be scoped, short-lived where possible and tied to a clear workload identity. Low-risk actions may be safe to automate. Destructive, irreversible or externally visible actions may require explicit authorization.

Runtime observability is just as important. Logging the initial prompt and final response will not explain a long-running agent's behavior. Defenders need visibility into tool calls, data access, memory changes, external connections and the sequence in which actions occurred. They also need the ability to interrupt execution, revoke access and reconstruct what happened afterward. MITRE's Adversarial Threat Landscape for Artificial-Intelligence Systems (ATLAS) now includes agentic AI in its threat matrix, giving defenders a shared language for this expanding attack surface.

Containment has to be designed before the incident. That means outbound network controls, tool-specific permissions, bounded execution time, resource limits and a reliable way to stop the agent without waiting for the model to cooperate. Sandboxes remain important, but their boundary must include every reachable tool and service. A tightly isolated model connected to an overprivileged tool is not a tightly isolated system.

Testing must expand from the model to the deployed system. The agent should be evaluated with the actual tools, permissions, memory, retrieval sources, MCP connections and approval workflows it will use in production. The National Institute of Standards and Technology's (NIST) AI Metrology Center provides a useful frame for selecting testing, evaluation, validation and verification methods across the AI lifecycle.

Finally, defenders need to decide where humans belong in the operating model. Requiring approval for every action erases the speed advantage that made the agent useful. Removing people entirely creates unacceptable operational risk. The practical design is governed autonomy: automate actions whose impact is understood and reversible, elevate consequential decisions and make every action observable and attributable.

The human role does not disappear. It moves from clicking through every step to defining boundaries, supervising outcomes and intervening when the system approaches those boundaries.

The real lesson from Las Vegas

The biggest lesson from Black Hat and DEF CON was not that AI is changing cybersecurity. Everyone already knows that.

The more important lesson was that autonomous agents are moving from research demonstrations into operating environments, security products and accessible offensive experimentation. Prompt injection, jailbreaks, model misuse and responsible AI remain important, but they now sit inside a broader problem: governing systems that can pursue goals and take action.

The AISI incident made it impossible to keep that argument entirely in the future tense. It did not prove that commercially available agents are escaping containment or attacking the internet on their own. It did show that capable agents given broad access and permissive controls may choose consequential methods their evaluators did not request or anticipate.

The organizations best prepared for that future will not necessarily have the most capable model. They will be the ones that can establish identity, constrain access, validate actions, observe runtime behavior and contain failures. They will also need to operate quickly because, as I argued in The model isn't the story. The response time is, capability matters only in the context of how fast defenders can respond.

Black Hat showed me what vendors are building to manage that future.

DEF CON showed me why they are building it.

For the first time, those two conversations felt perfectly aligned.