Blackhat and DEFCON34: Falling Back in Love With the Basics
In this blog
I spent the week listening for one thing, which was not what's coming but what we do about it.
Every answer I heard was twenty years old. Asset inventory. Identity management. Least privilege, configuration hygiene, testing, reduce your attack surface.
The panel of frontier lab security leads and former NSA and Cyber Command leadership closed by going around the table on the most important thing a defender can do, and the list was: leverage AI, asset inventory, testing, stop defaulting to a human in every loop, don't assume compliance will save you, build agility and stop assuming AI replaces humans. Build your bench.
That is my takeaway from Black Hat 2026. The most advanced security conference on earth spent a week telling us to do our homework.
Why that's the answer
"Do the basics" is easy to say and easy to wave away, so the reasoning matters more than the slogan does.
For years, roughly ninety zero-days were exploited in the wild annually. Discovery was hard and weaponization was harder, and that scarcity was doing far more work in our defensive model than most of us ever acknowledged. Our whole posture rested on it.
Because attacks were expensive, we got away with things. We left bitsadmin enabled across the estate because somebody, somewhere, might need it. We ran asset inventories eighteen months stale. We over-provisioned access because tightening it meant a fight with the business, and the fight was never worth having. None of that was ever fine, but it was survivable, because there were not enough cheap attacks in the world to find every gap we left open.
That has ended. CVEs are climbing and exploitation costs are falling. Models find flaws quickly, and the harness built around a model improves results even when the model itself is constrained. One of the sessions demonstrated the point live by directing an agent to build ransomware, and the striking part was not that it succeeded but that the model appeared to recognize what it was producing and produced it anyway.
The consequence is the thing I keep coming back to. Cheap offense did not invent new doors. It walks through the ones we left open. Configuration problems, rather than exotic vulnerabilities, account for ninety percent or more of what is genuinely exploitable.
So the basics did not come back. They never left. What we lost is the margin for error. Attack scarcity was quietly subsidizing two decades of not quite getting around to it, and that subsidy has now been withdrawn.
One of the lab security leads gave the week its best line: AI does not change the playbook, it added a fast-forward button to the playbook. Same plays, same techniques. Our playbook did not change either. What changed is that we no longer get a grace period.
The last basic
One more, from the last session I attended, which seemed unrelated to everything else all week.
It was about teenagers and about where the next generation of defenders comes from. The premise is that a meaningful share of kids commits something that would count as a serious cyber offense before they finish school, and that the natural talent pool is gamers and modders, because the skills transfer directly. This is not a fringe claim. UK National Crime Agency research has found that a large majority of young people recruited by online criminals developed their skills through gaming.
The case study was Marcus Hutchins. He started on hacker forums as a kid, wrote malware and was recruited into a criminal crew. He then stopped WannaCry, sparing a great many organizations serious damage, and it was that act that led investigators back to his earlier work. An FBI panelist described the same pattern from the other side of it: a fifteen-year-old who had compromised ships, teenagers who often cannot practically be charged, an attempted diversion program that didn't take.
The detail I have not stopped thinking about is that Hutchins did not know legitimate cyber jobs existed. World-class skills, and no awareness that there was a legal market for them.
Two days earlier, a panel of frontier labs and former NSA leadership closed with the instruction to stop assuming AI replaces humans and build your bench. Same conclusion, arrived at from opposite ends of the industry, one from juvenile cybercrime and one from frontier AI. Developing people is the most basic item on the entire list, and it's also the one we are most likely to believe AI exempts us from.
What I think we should do
Asset inventory first, because it surfaced in every session that mattered, it is the least interesting item on any list, and everything else depends on it. We cannot reduce a surface we cannot enumerate.
Then identity, which was called the single most important control by the person best placed to know, and nobody on the panel pushed back on her.
On least privilege, two competing product categories now exist to do this for us, and we should have a position on which one fits our environment rather than discovering our position in front of a customer.
And we should build the bench, since the people capable of operating any of this are frequently the same people who, at fifteen, looked like somebody's problem.
There was real fear in Las Vegas this year and some of it was earned, because models are getting out of their sandboxes and that isn't hype. But nobody on any stage suggested the answer was something we don't already know how to do. They said the answer is the thing we keep putting off.
Attackers change the economics and we can change the physics.
DEFCON34
Tony Martin and Arie Haenel, two Intel security researchers, spent their DEFCON talk making one point crystal clear: this is a system problem, not a prompt problem. A prompt is a request. It can't invent threat modeling context it was never given, and it can't manufacture the "build truth" that only comes from real source code and real specs. You can rewrite that prompt for hours and still be polishing the wrong tool.
Their answer wasn't a smarter prompt. It was a harness: a five-stage pipeline that treats an AI model less like an oracle and more like a very fast, very opinionated junior analyst who needs supervision, structure and a healthy dose of skepticism.
The wrong engine for the job
Before getting into the stages, the talk made a simple but easy-to-forget point: not every task needs your most powerful model.
Their analogy: you don't bolt a tank engine onto a shopping cart just to go buy milk. A scooter engine gets you to the store fine. Save the tank for the job that actually requires it.
In practice, that breaks down into three tiers:
T1 (light) – routing, dispatching, structured busywork
T2 (mid) – reading docs, building threat-model context, prepping the ground
T3 (frontier/"Mythos") – deep vulnerability reasoning, but only on a slice you've already narrowed down
The five stages: Every claim must earn its way forward
The core of the harness is a pipeline where nothing gets trusted until it's proven. Think of it less like a single conversation with an AI and more like a courtroom process: a claim enters as a rumor and has to survive cross-examination before it's allowed to become a finding.
Stage 1 (Scope & Compress): Define the Mission First
Before any hunting happens, the system scopes the target and compresses the corpus (documentation, source, build data) into something the model can actually reason over safely and accurately.
A great detail from this stage: documentation is a precision control, not a courtesy. An "obvious" bug in code might actually be an intended constraint documented in a spec the model never saw. The talk's example: a tax calculation that looks like a bug in isolation is actually correct behavior once you read the one line in the spec explaining the tiered rate. Skip the docs, and your model becomes an overconfident intern flagging things that were never broken.
To find the right passage instead of drowning in PDFs, Word docs and spreadsheets, they use a retrieval pipeline (lexical search plus vector search, fused and ranked) so the model gets a grounded quote with a real citation (document, line number) instead of a vague paraphrase.
Stage 2 (Hunt & Route): Chase Attack Paths, Not Files
Instead of scanning every function line-by-line, the system hunts attack paths, parallelized by attack surface. Routine audit categories (like running something through a standard driver security checklist) get handed to smaller, cheaper models. The frontier reasoning is reserved for the narrowed slice that actually deserves it.
The output of this stage isn't a verdict; it's a candidate. Every finding has to carry its location, evidence (an exact quote, not a summary), analysis, exploit steps and provenance. As the presenters put it plainly: severity and fixes wait until the claim survives triage.
Stage 3 (Dedup): One Root Cause, One Evidence Packet
Multiple hunts, multiple models, multiple passes: you end up with a pile of overlapping findings. Stage 3 consolidates them down to one root cause with one evidence packet, tagged with the exact build, module and offset.
Stage 4 (Triage & Prove): Agreement Is Not Proof
This might be the sharpest insight in the whole talk. It's tempting to think that if a second LLM reviews a finding and agrees, you've got independent validation. You don't. If the verifier is handed the same report, the same selected evidence and the same hidden assumptions as the hunter, all you've proven is that two models can nod along to the same story. Agreement is not proof.
Real verification means re-opening the primary sources (the actual source code, the actual threat model, the actual build) and re-grounding the claim from scratch. Fresh context can reduce anchoring bias, but it cannot replace re-grounding in reality. Every claim gets one of three outcomes: reject, downgrade or advance. Notably, the team found that source-grounded review removed far more false positives than review based on the report alone: proof that where you point the second look matters as much as that you took one.
Stage 5 (Verify & Learn): Reality Gets a Vote
Even a well-argued finding is still a hypothesis until it's tested against something real. Stage 5 runs the claim through a sandboxed "reality rig" (no secrets, no credentials, brokered egress only) and checks whether the proof-of-concept actually fires.
A failed proof-of-concept does not equal a false claim. Something can be real and still not reproduce cleanly on the first try, because of environment quirks, missing preconditions or a hardware clamp nobody documented. The point of this stage isn't just pass/fail; it's building verified knowledge: confirmed reachability, dead paths, hardware constraints, missing entry points. That knowledge feeds directly back into Stage 1 for the next scan, so the evidence packet starts stronger every time.
As the talk put it: the system learns through artifacts, not model memory. The model doesn't "remember" your codebase between sessions; your harness does, by getting smarter documentation into Stage 1 next time.
Five things to take home
The presenters closed with a simple checklist worth putting on a whiteboard:
Scope: Map the threat before you scan anything
Route: Send work to the right-sized engine, not the biggest one
Ground: Stamp every claim to its source
Challenge: Expose contradictions before they become reports
Learn: Feed verified evidence back into the next run
The specific use case here is vulnerability hunting, but the pattern generalizes to nearly any high-stakes AI workflow your organization is building: customer-facing agents, financial analysis, compliance review. Anywhere a model's output turns into a decision, the same discipline applies: scope it, route it to the right tier, ground it in evidence, challenge it independently and verify it against reality before anyone acts on it.
The uncomfortable truth is that a frontier model dropped into a vague prompt will happily generate something plausible. Plausible isn't the bar. Trustworthy is. And trustworthiness isn't a property of the model; it's a property of the system you build around it.