There's a reasonable-sounding argument moving through engineering organizations right now: AI plans the work. The agent writes the code, runs it, sees it fail and tries again until it passes. Brute force gets us to a working outcome. 

So, why spend senior hours on test-first discipline? Test-driven development (TDD) is a workaround for human fallibility, and the machine doesn't have that problem.

Part of that argument is right. The part that's wrong is expensive, and the bill arrives about two quarters late. The trouble is that brute force needs something to push against.

The loop must know when to stop

An agent works by trying something, checking the result and trying again. That checking step is the whole game. Without it, there is no loop at all, just an expensive guessing game with confident narration.

That's why AI has progressed further on code than most other work we've handed it. Code arrives with fast, cheap, impartial checks already attached: compilers, type checkers, linters, tests. Ask an agent to write a go-to-market strategy and nothing in the task itself can tell it when to stop. Ask it to turn a test suite green and it has a target it can measure itself against. That's the real advantage, and it's also where the trouble starts, because an agent pointed at a target will find a way to hit it.

Tests in an AI-native SDLC have stopped being quality assurance. They have become the finish line that tells the agent it is done. Take them away and you haven't cut ceremony; you've cut the mechanism that makes the agent work.

That also changes who the tests are written for. They used to serve the next engineer to touch the code. Now they serve the machine doing the work today.

What actually dies

The instinct that something has changed is correct, so let's be precise about what. Test-driven development does three jobs:

  1. It pressures your design. Writing the test first makes you use your own interface before you build it.
  2. It keeps the problem small. Tiny steps fit inside one person's head.
  3. It catches regressions. It serves as a net under you when you change something later.

The first two jobs are workarounds for human limits and don't transfer to an agent. Thoughtworks tested this directly by running test-first coding agents against agents without tests. The test-first runs didn't produce better protection against bugs. In some cases, their designs were worse because working in tiny steps suppressed the up-front thinking that drives agent quality. And they burned three to eight times more tokens.

The third job of catching regressions survives and grows. While handwriting tests first is losing justification, having the tests is becoming load-bearing. The underlying principle, that tests constrain what gets generated, holds up well. Give a model the tests along with the problem and its output measurably improves.

Two ways a green test suite lies to you, and both survive code review.

The first is self-marking. The model misreads a requirement, writes the code, then writes tests for the code it just wrote. They pass. They will always pass, because both came from the same misunderstanding. One mind created the exam and then took it. Green suite, wrong product. Worse, that green suite is now more dangerous than no suite at all, because it manufactures confidence. About one in 10 teams report handing back testing to the same AI that wrote the feature.

The second is what happens when the test becomes the target. Any measure you optimize hard enough stops measuring what you wanted. Researchers have caught coding agents doing exactly this: Building a fake path through the code that exists only to satisfy the assertion, while the feature you asked for never gets written. The agent delivers what you check, not what you requested.

That constrains the fix. The usual answer is to test your tests: deliberately introduce bugs and confirm the suite catches them. Good practice, and you should use it. But it only proves your tests can spot a broken implementation. It can't tell you the implementation is a decoy. It's also slow and compute-hungry, which is why it belongs on the code where failure is expensive rather than everywhere.

The only durable defense is structural: What "correct" means must come from outside the AI writing the code. A person writes it, a stakeholder signs off on it or a separate process derives it. It does not come from the same pass that produced the implementation. Running a second agent on the same prompt is not separation, because a misread requirement is still misread the second time.

What behavior-driven development (BDD) is actually for now

If TDD's role narrows, behavior-driven development's role grows, though not for the reason usually given.

BDD's value was never the given/when/then format. It was writing down what the system should do in language that a non-engineer can check. In a traditional software development lifecycle (SDLC), that was overhead; someone had to translate it into code anyway. In an AI-native SDLC, the same document does three jobs at once: it states the requirement, it instructs the agent and it defines what done looks like. One artifact, three uses, nothing lost in translation.

That's the case for specification-driven development, which means writing the behavior specifications first and letting agents build from them. But do it carefully, because the backlash is already underway and your engineers have heard about it. Critics call it "waterfall in new clothing." The sharpest version of the objection: a heavy up-front specification assumes you'll learn nothing useful while building, which is the assumption that agile existed to kill. The tooling is young and thinly adopted. 

The durable insight here has nothing to do with writing bigger specifications. It is the same separation principle as above, moved up a level, from "what should this function return" to "what are we actually building." And it points somewhere counterintuitive: Skipping this on new projects is backwards. New projects are where intent is least written down, least shared and most likely to drift. That's exactly when an agreed definition of done pays for itself.

Legacy systems: The real argument is throughput

The case for testing existing systems is usually made around safety. An agent will confidently rewrite code whose purpose it doesn't understand, and the tests are often the only place that purpose was ever written down. True, and reason enough on its own. But there is a sharper argument, and it is the one that moves budgets.

The evidence on AI-assisted development is contradictory. Controlled studies report productivity gains from 21% on enterprise tasks at Google to 56% on well-scoped benchmarks. Yet, a randomized METR trial found experienced developers were 19% slower with AI on mature codebases. The developers in the trial had forecasted a 24% speedup, and even after finishing the work, they still believed they had been 20% faster. 

The caveats matter: 16 developers, 246 tasks, early-2025 tooling. It is important to explicitly state that METR now treats the result as historical and believes it is likely that developers are faster with AI tools now. What still holds up in 2026 is the gap between what the developers felt and what the clock said, and that gap is what engineering leaders keep describing.

Broader telemetry points the same way. Faros 2025 research showed that across more than 10,000 developers, merged pull requests are up 98%, review time is up 91% and average change size is up 154%, while delivery metrics at the company-level stay flat. A year of further data from Faros suggests the bottleneck has not eased so much as moved: merged pull requests with no review were up 31% in spring 2026 and climbed to 76% in fall 2026. Time spent in QA is up 301%. DORA research associated a 25% rise in AI adoption with a 7.2% drop in delivery stability. 

Three things explain the split: how abstract the task is, how experienced the developer is and how mature the codebase is. That last one matters most here. New code benefits enormously, while mature codebases pick up verification overhead that can swallow the generation savings whole. The business case follows from that.

On your most valuable systems, AI is producing code faster than your organization can convince itself that the code is safe.

The gain is real. It's being eaten by the review queue before it reaches a customer.

Tests are how you attack that queue, and the reason to fund them is throughput rather than insurance. A good suite turns "two senior engineers read this change for a day and a half" into "the pipeline answered in four minutes." That's a cycle-time argument, and cycle time is a number your CFO already tracks.

Two honest caveats for engineering leadership

First, tests collapse correctness checking, not design review. With change sizes up 154%, a lot of that review time is senior people asking whether a sweeping change is a good idea, which no test can answer. Second, retrofitting tests onto a large legacy system is a multi-quarter investment with real cost. Price it. The argument is that the payback lands in delivery speed rather than defect counts, and that it's currently the binding constraint on getting any AI value out of the systems you care about most. The claim is defensible, but it is not free.

New code goes stale faster than your policy assumes and that undercuts the tidy new-versus-legacy split this article leans on. That split assumes a stretch of time where code is new and the team knows it intimately. AI compresses that from both ends: volume per week is up sharply, familiarity per line is down just as sharply, and churn is rising. 

Your new project becomes a mature codebase fast, and it arrives there as genuine legacy: code running in production that nobody on the team has read. The safety net you skipped at the start goes missing well before anyone declares a maintenance phase.

Replace the binary with three questions

New versus legacy is a useful instinct and one axis of a three-axis decision. Scale the rigor by scope, maturity and consequence:

SituationWhat to require
Small change, new code, cosmetic riskNothing formal. Generate it and read it. This is where the big speed gains are real.
Small change, mature code, functional riskA plain-language description of expected behavior, plus targeted tests, before you accept the output.
Sweeping change, functional riskA written spec, approved before generation starts.
Auth, payments, personal data, data integrityHard rules encoded once and enforced on every generation.

 

The key takeaway is that the requirements change from task to task inside a single sprint. Blanket mandates fail in both directions: they tax trivial work into paralysis and wave through the change that takes down production.

What to tell your teams

  • Own the definition of done; let the machine write the tests. Senior effort moves from writing test cases to deciding what correct means. For most senior engineers, that is a promotion.
  • Never let one pass produce both the code and the standard it's judged against. The standard comes from outside, every time.
  • Assume the agent will build to the test. Read the implementation, not just the result. A passing suite is evidence rather than proof.
  • Fund test coverage on your highest-value legacy systems. Price it honestly as a speed investment, because that is where AI is underperforming today.
  • Match rigor to risk, task by task. A single policy applied across a whole team or a whole quarter will be wrong in both directions.

TDD and BDD were never ceremonies that we tolerated because humans are unreliable. They were the practice of saying what we wanted precisely enough to check it. We've just handed the building to something fast, tireless and completely unable to guess what we meant. Saying what we want precisely has never mattered more.

At World Wide Technology, we help enterprises build the testing, governance and delivery practices that turn AI-native engineering from individual speed into organizational throughput, and we prove them in our AI Proving Ground before they touch production. 

Ready to talk with our AI-native engineering team for real-world guidance? Request briefing