TL;DR: Frontier AI models now refuse the exact work defenders do (analyzing exploits, malware and live-incident artifacts) while attackers, running jailbroken or open-weight models under no usage policy, face no such limits. Some now lace malware with content engineered to trigger those refusals. WWT recommends that organizations start discussions to decide how they can provide operational agility to stop operating at the default refusal boundary. 

Apply to Anthropic's Cyber Verification Program  and OpenAI's Trusted Access for Cyber. Both programs lower the refusal threshold for authorized defensive work such as vulnerability and exploitability analysis, malware and binary reverse engineering, detection engineering, patch validation, whilst keeping genuinely prohibited categories such as ransomware, mass exfiltration, and command-and-control blocked for everyone. Vetting is fast; some applicants report approval within a business day. Every company that defends software should have at least one verified-defender account in place before the next incident, not during it. Individuals verify at https://chatgpt.com/cyber; enterprises apply through OpenAI's TAC form  and Anthropic's CVP.


In June I wrote that the hard part of cybersecurity had moved. Frontier AI had made finding vulnerabilities cheap and fast, and the real contest was remediation velocity — the distance between discovery and a working fix in production.

That is still true. But the Hugging Face intrusion disclosed on July 16 revealed a second asymmetry sitting underneath the first, and it is the one that should worry security leaders most. It is not about who can find bugs faster. It is about who is allowed to act. In that incident, the attacker operated under no usage policy at all, while Hugging Face's own incident responders were locked out of their best tools at the worst possible moment — not by the attacker, but by our own industry's safety controls.

An autonomous AI agent system breached part of Hugging Face's production infrastructure. It entered through the data pipeline via a malicious dataset that abused two code-execution paths and then escalated to node-level access, harvested credentials, and moved laterally across internal clusters over a single weekend, generating more than 17,000 recorded actions from a swarm of short-lived sandboxes with self-migrating command-and-control. This is the agentic attacker the industry has been forecasting, now confirmed in production.

To reconstruct what tens of thousands of automated actions had done, Hugging Face's responders reached for frontier models behind commercial APIs. It did not work: the analysis required submitting real attack commands, exploit payloads, and command-and-control artifacts, and those requests were blocked by safety guardrails that could not tell an incident responder from an attacker. The team pivoted to an open-weight model — GLM 5.2 — on their own infrastructure. It worked, it kept stolen credentials inside their environment, and it turned days of work into hours. Hugging Face called the problem "guardrail lockout," and recommended every defender keep a capable, self-hosted model vetted and ready before an incident.

The most capable defensive tooling available was withheld from the defender mid-breach because the payload looked dangerous. It looked dangerous because it was dangerous — that is the nature of the work.

The logic of a refusal is that the requester might be an attacker. But the attacker in this incident never touched a guardrailed commercial endpoint. Attackers use jailbroken or open-weight models, or models with the safety weights stripped out — a process now cheap enough that researchers have removed guardrails from mainstream open models in minutes. The guardrail was never going to see them. It falls instead on the party guaranteed to be inside the policy: the defender, the researcher, the responder, the SOC that pays for an API and agrees to its terms. We have built a safety mechanism whose burden lands on the compliant while its intended target routes around it. The people who follow the rules are the only ones the rules can reach.

There is a sharper turn. If a refusal reliably blocks a defender's AI, the refusal itself becomes a control the attacker can trigger on purpose.

This is already happening. In 2025, Check Point Research documented the first malware sample carrying an embedded prompt injection in which a plain-language instruction was embedded telling any analyzing model to ignore its task and call the file benign. It read like a proof of concept; it did not stay one.

The tactic has matured into a deliberate evasion class. Researchers now describe malware seeded with "refusal-bait": blocks of prohibited content (chemical, biological, radiological, and nuclear (CBRN) weapon-making text, exploit-development material, and similar tripwires) inserted for no functional reason except to make an LLM-based scanner hit its own safety filter and refuse to analyze the file. The refusal meant to prevent harm becomes the thing that lets the malware pass unread. Analysts tracking these families increasingly treat embedded safety-trigger strings as standard tradecraft, alongside packing and obfuscation.

Cloudflare's Cloudforce One measured the effect: with "Notice to AI" lures and payloads buried in large code bundles, detection accuracy fell to as low as 12%. Palo Alto Networks' Unit 42 now lists "AI access restriction", tripping a model's safety filters so it refuses to process content, as a named technique in its in-the-wild taxonomy.

The two failure modes compound: the attacker plants a tripwire to force a refusal, and the defender's own model refuses because the material looks dangerous which it always was, because analyzing danger is the job. A control meant to prevent harm can end up shielding the attacker.

None of this is an edge case. Security analysts flagged the broader pattern months ago: in March, CSO Online documented that AI guardrails increasingly block legitimate security work while attackers bypass them with ease, leaving defenders unable even to test their own exposure. In practice, these guardrails often function more like negotiable social norms than hard boundaries.

This is not an argument against safety. The misuse risk is real: Anthropic disrupted an AI-orchestrated espionage campaign in late 2025, and autonomous tooling lowers the cost of patient, multi-stage attacks run at machine speed. Nobody seriously wants frontier cyber capability handed to anyone who asks. The problem is the instrument, not the intent. A blanket refusal on hosted models cannot distinguish a responder from an adversary, so it treats everyone as a suspect and obstructs the people it can reach while handing attackers a tripwire they can pull at will. Safety that only constrains the compliant does little to make anyone safer, and its cost shows up as longer dwell time.

There is a better idea on the table. Researchers at the Institute for AI Policy and Strategy have spent a year developing "differential access" — tilting the balance toward defense by shaping who gets cyber-capable models rather than throttling everyone. Even in its most restrictive posture, the principle holds: give vetted defenders an advantage. Anthropic's Cyber Verification Program and OpenAI's Trusted Access for Cyber are early moves in that direction.

What security leaders should do now

  • Apply to the labs' defender programs today. Anthropic's CVP and OpenAI's TAC lower the refusal threshold for authorized defensive work; individuals verify at chatgpt.com/cyber, enterprises apply through the TAC form and CVP. Vetting can take as little as a business day — there is no reason to wait for an incident to begin.
  • Stand up a capable model you can run yourself, before an incident, not during one. It removes the risk of guardrail lockout, keeps attacker data and stolen credentials inside your environment, and lets you tune out refusal-bait instead of being blocked by it. Your model choice will change every quarter; the pipeline and the muscle memory of running forensics on infrastructure you own will compound for years.
  • Don't rely on a single AI verdict. Pair model-based analysis with deterministic tooling — strings, YARA, static analysis, isolated sandboxes — so a forced refusal or a poisoned classification is never a single point of failure.

What providers and policymakers must change

  • Build a verified-defender lane. Vetted teams, responders, and researchers should reach full-capability models through credentialed access, not the same refusal path as an anonymous prompt.
  • Recognize incident-response context. A request that submits exploit payloads and command-and-control artifacts for analysis is the signature of defense, not offense and embedded refusal-bait should not be allowed to dictate the verdict.
  • Publish the carve-out, and measure the friction. Document plainly what a model will and won't do in an incident, and track false positives on legitimate security work as the risk transfer they are, then drive them down.
  • Enable defenders; don't only restrict them. Export-control and licensing reflexes strip capability from defenders first. The White House AI Action Plan already urges critical infrastructure to adopt AI for defense and governments should fund that, vet defenders, and scale credentialed access well beyond today's handful of partners.

WWT's point of view

We saw this coming. Our customer conversations signaled the shift toward agentic attacks in early 2026, and we invested ahead of it. One such example is building new roles on our Global Cyber team, such as Defensive AI Security Scientists, to get ahead of the threat and help customers reach a real state of readiness. A year ago, that role did not exist. That is what preparedness looks like in practice: capability, process, and people in place before the incident, not scrambled together during it.

Our view is straightforward. This asymmetry is policy-made, which means it is reversible. The capability question is settled — machines now find bugs, run campaigns, and reconstruct attacks faster than people can. The open question is whether we let defenders use that capability or keep handing the advantage to attackers by default. Right now the attacker has no guardrails and the defender does. That is not a law of nature; it is a policy we wrote, and we can rewrite it by giving vetted defenders an asymmetric advantage instead of an asymmetric handicap. The attacker already runs unrestricted. It is time we stopped restricting the defender.