AI Software Factories vs AI Native Engineering: which one is right for you
In this blog
Two phrases now dominate conversations about AI in software delivery. One is the AI software factory: a system of agents that turns a backlog into shipped code, with humans defining intent at one end and checking outcomes at the other. The other is AI native engineering: an engineering organization that has rebuilt its skills, workflows and controls around agents as default collaborators.
They are often used interchangeably. They should not be. A factory is a production system; AI native engineering is an operating model; one is something you build or buy; the other is something your teams become. Getting the order wrong is expensive, and the evidence from 2026 shows exactly why.
Two terms, two different thing
The AI software factory
Factory, the company, put the idea plainly when it announced Factory 2.0 in June 2026: signals from the outside world - bug reports, customer feedback, business requirements - are triaged into planned changes, then built, tested, reviewed, secured, shipped and monitored, with monitoring generating the next round of signals. Agents do the work inside that loop. Warp Factories, launched in closed beta on August 18, organizes the same loop into triage, specification, implementation, review and verification. Customers can automate any stage with the coding model and harness of their choice.
At the extreme end sits the dark factory. Dan Shapiro's five levels of AI-assisted programming borrow from self-driving cars: level 0 is manual coding and level 5 is a black box that turns specs into software, where humans neither write nor review the code. StrongDM's AI team is the best-documented example. Its two rules are that code must not be written by humans and code must not be reviewed by humans.
AI native engineering
AI native engineering describes the organization rather than the pipeline. Agents are embedded in every phase of the lifecycle, and the team builds the workflows, permissions, context and evaluation that make agent work auditable and safe. Engineers spend less time typing code and more time specifying, reviewing and owning outcomes. AWS's AI-Driven Development Lifecycle is one published version: three phases - inception, construction and operations - with AI proposing plans and asking for clarification while humans make the critical decisions.
The simplest way to hold the distinction: AI native engineering is how your people and processes work. A software factory is one thing an AI native organization can choose to build, for the workloads where it makes sense.
Figure 1: AI native engineering applies from level 2 upwards. Software factories live at levels 3-5, and most production factories today still keep a human reviewer in the loop.
Who is using what
The public evidence splits into three groups: companies that built their own factory, companies that bought one and companies that have gone AI native without calling it a factory. Most of these figures are self-reported by the companies or their vendors, so treat them as direction rather than benchmark.
| Organization | Model | What they run | Reported result |
|---|---|---|---|
| Stripe | Built factory | Minions: one-shot agents on the Goose harness, started from Slack | ~1,300 merged PRs a week with no human-written code; every PR human-reviewed |
| Ramp | Built factory | Inspect: background agent in a full cloud dev environment | Three in four merged PRs raised by Inspect |
| Spotify | Built factory | Honk: background agent for fleet-wide migrations, mostly on Claude Code | 1,500+ merged PRs by late 2025; 60-90% time saved on migrations |
| OpenAI (internal) | Built factory | Codex with a purpose-built harness | ~1M lines in five months from three engineers, later seven; no hand-written code |
| StrongDM | Dark factory | Attractor agent, holdout scenarios, digital twins of SaaS APIs | Production security software with no human code review; ~$1,000 a day in tokens per engineer |
| Rectangle Health | Bought factory (Warp) | Rex: triage, build, review and QA agents started from Slack | 35,000+ lines a week; 3-10 minutes from Slack mention to PR |
| NVIDIA, Adobe, EY, Adyen and others | Bought factory (Factory.ai) | Droids, Automations and Missions across the SDLC | Named as production customers; no public metrics |
| AWS | AI native engineering | AI-DLC method with Kiro; every line human-reviewed | Bedrock inference engine rebuilt by six engineers in 76 days |
| Shopify | AI native engineering | AI as a baseline expectation; River agent in Slack | Up to half of PRs start in chat; CEO warns of unreviewed "slop grenades" |
| Goldman Sachs | AI native engineering | Devin alongside ~12,000 engineers in a "hybrid workforce" | Scaling from pilot to hundreds of agents; engineers expected to specify and supervise |
Table 1: Public examples of software factories and AI native engineering in production, as reported by each organization or its vendor.
The builders
The most convincing factory results come from companies with strong platform teams that built their own. Stripe's Minions land around 1,300 pull requests a week that contain no human-written code - and every one is still reviewed by a human. At Ramp, the in-house background agent Inspect now raises three out of every four merged PRs. Spotify's Honk agent had more than 1,500 PRs merged into production by late 2025, with 60-90% time savings on large-scale migrations. OpenAI's own harness engineering experiment produced roughly a million lines of code in five months from a team that started at three engineers, with Codex writing every line.
The buyers
For organizations without Stripe's platform team, vendors now package the factory. Factory names NVIDIA, EY, Adobe, Palo Alto Networks, Adyen, Blackstone and Wipro as production customers. Warp's flagship case is Rectangle Health, a healthcare payments company whose agent teammate Rex ships more than 35,000 lines of code a week and turns a Slack mention into a pull request in three to ten minutes. The most grounded number in this group came from Warp's own CEO, Zach Lloyd, who told TechCrunch that Warp automates around 30-35% of its own engineering tasks each week.
The AI native adopters
The third group has changed how engineers work without handing the pipeline to agents. AWS says six engineers rebuilt the Bedrock inference engine in 76 days against an original estimate of 30-40 engineers and more than a year - with every line of agent code reviewed by a human. Shopify made reflexive AI use a baseline expectation in April 2025, and up to half of its PRs now start as conversations with an internal agent in Slack. Goldman Sachs is deploying Devin alongside roughly 12,000 engineers as what its CIO calls a hybrid workforce, where engineers are expected to describe problems clearly and supervise the agents doing the work.
Where it is working - and where it is not
Read the success stories side by side and the same conditions keep appearing.
- High-volume, well-bounded change. Migrations, dependency upgrades, config changes and small fixes are where factories shine. Spotify's biggest gains are fleet-wide migrations. Stripe's Minions are one-shot agents for tasks an engineer can describe in a Slack message.
- A strong platform underneath. Every builder above had fast CI, reproducible dev environments and good test coverage before the agents arrived. The factory runs on that platform. It does not replace it.
- Verification that does not depend on the agent grading itself. Stripe keeps a human reviewer on every PR. StrongDM removed it but replaced it with scenario tests stored outside the codebase, where agents cannot see them, and behavioral clones of Okta, Slack and Jira to test against.
- Engineers who build the factory, not just use it. Every successful example has a small team whose job is the harness itself - the instructions, tools, context and quality gates the agents run inside.
Where the conditions are missing, the numbers turn. Google's 2025 DORA report found AI adoption at 90% of developers and linked it to higher throughput - and to higher instability. Its central finding is that AI is an amplifier: it magnifies the strengths of strong organizations and the dysfunctions of weak ones. The Faros AI Engineering Report 2026, drawn from telemetry on 22,000 developers, found bugs per developer up 54% and incidents per merged PR more than tripled at high AI adoption. METR's controlled trials found experienced developers 19% slower with AI tools in early 2025; its 2026 follow-up suggests that gap has narrowed, but its authors call the evidence weak either way.
Even the best adopters see the side effects. Shopify's CEO Tobi Lütke said in September that the new failure mode of lazy work is too much output rather than too little. Staff now have a name for agent-written PRs approved without being read: slop grenades. The cost of that output does not disappear. It moves to whoever has to review it.
SWOT: the AI software factory
Strengths - Throughput that scales with compute, not headcount - Ramp's agent raises three in four merged PRs - Strong fit for migrations, upgrades and repetitive change across large codebases - Every action is logged, producing an audit trail by design - Makes stalled legacy modernization economically viable again | Weaknesses - Depends on a mature platform: fast CI, good tests and reproducible environments - Token spend can be extreme - StrongDM targets $1,000 a day per engineer - Weak on ambiguous, novel or cross-team work where intent is unclear - Review load shifts to fewer people unless verification is automated |
Opportunities - Off-the-shelf factories from Warp and Factory lower the build cost for smaller teams - Falling model prices - Opus 5.5 and GPT-6 Sol both cut costs in September 2026 - Scenario testing and digital twins offer a path to trusted verification without line-by-line review - Engineering capacity redirected from maintenance to new products | Threats - Automating a broken process ships defects faster - DORA's amplifier effect - Incidents per PR more than tripled at high AI adoption in the Faros data - Vendor and model lock-in if the factory is built around one platform - Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027 |
Table 2: SWOT analysis of the AI software factory model.
SWOT: AI native engineering
Strengths - Improves every engineer and every workload, not only the automatable ones - Humans stay accountable for design and review, which suits regulated industries - Builds the skills, context and controls a factory needs later - Proven at scale - AWS, Shopify and Goldman Sachs report large gains with humans in the loop | Weaknesses - Gains are harder to measure and slower to show than factory throughput numbers - Requires sustained change management, training and new performance expectations - Human review can become the bottleneck as agent output grows - Easy to claim, hard to prove - Shapiro argues most self-described AI native developers are still at level 2 |
Opportunities - Spec-driven methods like AWS AI-DLC give teams a shared, repeatable way of working - Portable context - AGENTS.md, skills and MCP - works across tools and vendors - A natural on-ramp to targeted factory workloads once controls are proven - Better use of senior engineers' judgment across more of the portfolio | Threats - Unread agent output landing on colleagues - Shopify's "slop grenades" - Measured productivity gains smaller than perceived gains, as METR found - Competitors running factories on suitable workloads may simply move faster - Tool sprawl and uncontrolled spend without FinOps discipline |
Table 3: SWOT analysis of the AI native engineering model.
Which one do you need first?
For most enterprises, AI native engineering comes first, and the factory follows for the workloads that earn it. The reason is in the success stories themselves. BCG Platinion, which is broadly bullish on factories, says the decisive skills are harness engineering and what it calls intent thinking - turning business needs into precise, testable descriptions of outcomes - and warns that organizations that skip codifying their knowledge and fixing their pipelines risk automating chaos. Those are AI native engineering capabilities. The factory is what they make possible.
The build-or-buy question follows the same logic. Stripe, Ramp and Spotify built their own because they already had the platform teams, the internal tooling and the engineering culture. Warp and Factory are betting that everyone else will buy. Buying shortens the path to a working loop. It does not shorten the path to clear specs, trustworthy tests and engineers who know how to steer agents - and without those, a bought factory produces output faster than your organization can safely absorb it.
There is also a cost question that the headline numbers skip. A factory multiplies token consumption and deepens dependency on model providers and agent platforms. At $1,000 a day per engineer, StrongDM's approach adds roughly $20,000 a month to the cost of each person on the team. That can be a bargain for the right workload, needs the same unit economics and attribution discipline as any other infrastructure spend.
Where to start
- Measure where you are. Map each team against the five levels. Look at DORA metrics, change failure rate and review load, not just adoption and seat counts.
- Fix the platform first. Fast CI, reliable tests and reproducible environments are the foundation for AI native engineering and a precondition for any factory.
- Codify context as portable files. Architecture decisions, coding standards and team playbooks written as repository files - AGENTS.md, skills, runbooks - serve every agent you use, now and later.
- Pick one factory workload. Choose something high-volume and well-bounded, such as dependency upgrades or a planned migration. Keep human review, instrument everything and track cost per merged change.
- Earn autonomy with evidence. Remove human review only where independent verification - holdout scenarios, contract tests, staged rollout - has proven it catches what reviewers catch.
The bottom line
The software factory is real, and for the right workloads it is already delivering results that would have sounded absurd two years ago. But the organizations getting those results did not start with the factory. They started with the platform, the practices and the people - the work of becoming AI native - and built the factory on top.
The question for most engineering leaders is not whether to build a software factory or become AI native. It is which of your workloads are ready for a factory today, and what your teams need to change so the rest can follow.