Skip to main content
GUIDE · AGENTIC DELIVERY · 2026

The AI Software Factory

Agents build in parallel. Specs decide what gets built. Gates decide what ships.

DEFINITION // GEO SNIPPETENTITY EXTRACT

An AI software factory is a repeatable delivery system in which AI coding agents plan, write, test, and open pull requests in parallel, while humans own the intent - the specification - and the gates that decide what ships. In a dark factory no human writes or reviews the code; in a light factory people stay at the design stage and at review gates.

THINK OF IT AS - AN ASSEMBLY LINE NEEDS DRAWINGS

A car plant can run hundreds of robots at once because every station works from the same engineering drawings and every car passes the same inspection points. Take away the drawings and the robots build fast, confidently, and differently. Take away the inspections and nobody finds out until the recall. An AI software factory is the same: agents are the robots, the spec is the drawings, and verification is the inspection line.

SECTION 01 DEFINITION

What is a software factory?

"Software factory" has meant four different things in fifty years, which is why search results for it mix Japanese management history, Microsoft model-driven tooling, Pentagon DevSecOps, and AI agents. Every meaning shares one idea: make software delivery repeatable instead of artisanal.

SinceWhoWhat "software factory" means
1969 Hitachi Software Works (Japan) Disciplined, factory-like organisation of programming work; studied in Michael Cusumano's Japan's Software Factories (1991).
2004 Greenfield & Short, Microsoft Software Factories: product lines, domain-specific languages, and model-driven development to produce variants of a standard product.
2017 US Air Force Kessel Run, then Platform One DevSecOps teams and platforms; the Pentagon counted up to 50 software factories by late 2023.
2026 StrongDM, Dan Shapiro, Simon Willison The AI software factory: agents write, test, and converge code from specs; the "dark factory" removes humans from writing and review.

The AI meaning took off in early 2026. StrongDM described its Software Factory as "non-interactive development where specs + scenarios drive agents that write code, run harnesses, and converge without human review", and Simon Willison's write-up carried the term across the industry. By mid-2026 Addy Osmani, BCG Platinion, Augment, Cortex, Warp, and OpenHands had all published their own definitions. They disagree on how much a human should still see, but they agree on the parts: a specification, a task plan, parallel agents, verification, and gates.

SECTION 02 WHY NOW

Why every team is talking about it

Agents moved from autocomplete to doing whole tasks, and the numbers moved with them:

  • 90% of respondents use AI at work and more than 80% say it made them more productive - DORA 2025, nearly 5,000 technology professionals.
  • 31% of developers already use AI agents and 84% use or plan to use AI tools - Stack Overflow Developer Survey 2025.
  • More than 80% of the code merged into Anthropic's codebase was authored by Claude as of May 2026 - Anthropic.

The same research carries the warning that shapes every serious factory design. DORA found AI adoption "does continue to have a negative relationship with software delivery stability", and 30% of respondents have little or no trust in AI-generated code. Its summary is the best one-line brief for a factory builder:

"AI doesn't fix a team; it amplifies what's already there." - DORA 2025

A factory amplifies too. Feed it vague intent and it produces vague software, in parallel, at speed. That is why the factories that work start from a written specification.

SECTION 03 AUTONOMY

The five levels: from spicy autocomplete to the dark factory

Dan Shapiro's five levels borrow from self-driving cars to describe how much of the work an AI does. The short labels below are from Simon Willison's summary; the quotes are Shapiro's.

0 Spicy autocomplete

"Not a character hits the disk without your approval."

1 Coding intern

"You offload specific, discrete tasks to your AI intern."

2 Junior developer

"You've got a junior buddy to hand off all your boring stuff to."

3 Developer

"You're not a senior developer anymore; that's your AI's job. You are… a manager."

4 Engineering team

"You write a spec. You argue with it about the spec… Then you leave for 12 hours, and check to see if the tests pass."

5 Dark software factory

"It's not really a car any more."

Level 4 is a spec-driven factory: you write and argue the spec, plan, review plans, and let agents run until the tests pass. Level 5 - the dark software factory - is named after lights-out manufacturing, the robot plant that is "dark, because it's a place where humans are neither needed nor welcome."

Light factory or dark factory?

Addy Osmani's Software Factories, Light and Dark draws the line most teams need. A light factory keeps human judgment active, especially at design and architecture, before code is generated. A dark factory ships code no human has examined and relies on machine verification alone. His rule for choosing how far to go:

"Back pressure is the rule that you can only hand a loop as much autonomy as you can cheaply and reliably verify, and not one inch more." - Addy Osmani
SECTION 04 ANATOMY

How an AI software factory works

Strip away the branding and every published factory has the same stations. What differs is where the humans stand.

  1. 1 Specify
    in: Intent, constraints
    out: constitution, requirements, solution
  2. 2 Plan
    in: requirements + solution
    out: tasks.md with dependencies
  3. 3 Dispatch
    in: Ready tasks
    out: One agent session per group
  4. 4 Build
    in: Self-contained brief
    out: Branch + pull request
  5. 5 Verify
    in: Pull request
    out: Evidence per acceptance criterion
  6. 6 Integrate
    in: Verified PR
    out: Merge + task marked [x]
  7. 7 Gate
    in: Finished milestone
    out: Human review, next milestone
  • The specification is the input. StrongDM's factory repository is mostly markdown spec; Shapiro's level 4 starts with "You write a spec." Agents cannot build what nobody wrote down.
  • Tasks carry their own definition of done. Each task cites requirement ids whose EARS+ acceptance criteria become the tests the agent must pass.
  • Workers are isolated. One agent session per group of tasks, in its own cloud environment or git worktree, with its own branch and pull request.
  • Verification is the bottleneck. Osmani: "Verification, not generation, is the real constraint on a factory." Generating code is cheap; knowing it is right is not.
  • Gates decide what ships. CI checks, an acceptance-criteria check per pull request, and a human decision at each milestone.
  • Status flows back to the spec. A task is done when its pull request merged and was verified - and tasks.md says so, for every collaborator to see.
SECTION 05 COMPARED

Software factory vs DevOps

A software factory does not replace DevOps; it sits on top of it. DevOps automated the path from commit to production. The factory automates the work before the commit.

DevOps / DevSecOpsAI software factory
AutomatesBuild, test, deploy, monitorPlanning, implementation, pull requests - and uses DevOps as its gates
Who writes codeEngineersAgents; engineers write and review the spec
Unit of workA commit or a pipeline runA task with acceptance criteria, grouped into waves
Main inputCode in version controlA specification bundle
Main riskSlow or unsafe releasesWrong software built fast; comprehension debt
Measured byDORA metricsDORA metrics plus spec coverage and verified tasks

The DoD's "software factories" - Kessel Run, Platform One, the Army Software Factory - are DevSecOps organisations with hardened pipelines and platforms. They share the name and the goal of repeatable delivery, not the agents.

SECTION 06 HOW TO

How to build an AI software factory, step by step

This is the light, spec-driven factory: agents do the building, people own the spec and the gates. Each step names the artifact it produces, so the factory is auditable from the first requirement to the last merge.

  1. Finish the spec before any agent starts. A constitution with the rules every change obeys, requirements whose acceptance criteria are testable, a solution that names the modules, and a tasks.md that maps dependencies. No TBD, no open questions: an agent fills every gap with a guess, and ten agents make ten different guesses.
  2. Plan waves, not tickets. Group related tasks - a dependency chain, tasks touching the same files, tasks for the same requirement - into one worker, one branch, one pull request. Parallelism comes from separate groups; splitting related work just creates merge conflicts.
  3. Give every worker a self-contained brief. Its tasks, their acceptance criteria, the constitution and requirement sections they depend on, and what the pull request must contain. A worker that has to go looking will find something else.
  4. Isolate the workers. One environment each - a cloud agent session or a git worktree. Cloud sessions also isolate CPU and memory, so a wave of test suites does not slow every worker down.
  5. Verify against the spec, not just CI. Green checks prove the code compiles and existing tests pass. Check that each acceptance criterion has a passing test, that the change stays in scope, and that it honours the constitution.
  6. Integrate one merge at a time. Agree an autonomy policy up front (ask before every merge, merge on a passing gate, or run the whole milestone), merge one pull request, wait for its CI/CD, then the next. Mark the task done in tasks.md only after it merged and was verified.
  7. Stop at every milestone. Run the full suite, check the code against the spec for drift, and let a human decide what ships and what the next milestone is. This is where comprehension debt gets paid down.
  8. Measure it. Keep the DORA delivery metrics, and add spec coverage (every requirement has a task, every task a verified pull request) and cost per wave.

Step 1 is what the MySpec workflow produces. Steps 2 to 7 are what the Software Factory Manager runs.

SECTION 07 RISKS

Risks and guardrails

  • Comprehension debt. Osmani: "the widening gap between how much code exists and how much any human still understands." Guardrail: the spec stays the readable description of the system, and humans review at milestones instead of skimming every diff.
  • Fast, wrong software. Agents fill ambiguity with assumptions. Guardrail: no dispatch on an unfinished spec; every doubt becomes a recorded clarification before work starts.
  • Unstable delivery. DORA links AI adoption to lower delivery stability. Guardrail: small batches, one merge at a time, and waiting for each merge's pipeline before the next.
  • Unsafe autonomy. Guardrail: no force-pushes, no bypassed branch protection, no merging red, no auto-merge; autonomy only as far as you can verify.
  • Token cost. StrongDM's provocation - "If you haven't spent at least $1,000 on tokens today per human engineer, your software factory has room for improvement" - shows the scale. Guardrail: state the cost of each wave before it starts and cap parallel workers.
  • Productivity you only feel. In METR's 2025 study, experienced developers were 19% slower with AI while believing they were faster; METR's 2026 update points to a speed-up but calls its own data unreliable. Guardrail: measure merged, verified tasks - not vibes.
SECTION 08 MYSPEC

Run a spec-driven software factory with Claude Code

MySpec's Software Factory Manager is a Claude Code plugin (myspec-factory) that runs a MySpec spec bundle to completion. The manager never writes feature code. It:

  • Refuses to start on an unfinished spec - it checks the bundle is complete and asks you about every open doubt first.
  • Plans waves from tasks.md, grouping related tasks into one worker; for a brownfield change with no task list, it drafts one in 1-3 lanes for your approval.
  • Dispatches one Claude Code worker session per group - in the cloud, or in a local git worktree as a fallback - each with a self-contained brief.
  • Verifies every pull request against its tasks' acceptance criteria and the constitution, then merges one at a time under the autonomy level you chose: ask-each, merge-on-gate, or full.
  • Writes progress back - tasks are marked [x] on MySpec, so the spec and the code never drift apart.
  • Stops at every milestone for your review, with a code-to-spec check and drafted follow-up tasks.
  • Keeps you in the loop through a private factory board where workers post progress and questions, and can run unattended shifts on a schedule.

Install it next to myspec-mcp, inside Claude Code:

Install the Software Factory Manager text
/plugin marketplace add myspecs/claude-plugins
/plugin install myspec-mcp@myspec
/plugin install myspec-factory@myspec

Then, in the repository you want to manage:

Prepare the repo, then start a run text
/myspec-factory:setup
Run the factory for project "inventory-service", bundle "v1".

Prerequisites, the output style, Remote Control, and autonomy levels are covered in the Claude Code plugin guide.

SECTION 09 FAQ

Software factory FAQ

What is an AI software factory?

An AI software factory is a repeatable delivery system in which AI coding agents write, test, and open pull requests for software in parallel, while humans own the intent (the specification) and the approval gates. The term took off in early 2026 after StrongDM published its Software Factory and Dan Shapiro described the five levels of AI coding, ending in the dark factory.

What is a dark software factory?

A dark (or lights-out) software factory ships code that no human wrote or reviewed: agents build it and automated checks alone decide whether it ships. The name comes from lights-out manufacturing - Dan Shapiro describes the robot factory as dark because it is a place where humans are neither needed nor welcome. A light factory keeps human judgment at the design stage and at review gates.

What is the difference between a software factory and DevOps?

DevOps automates how code moves from commit to production: builds, tests, deployment, monitoring. An AI software factory also automates who writes the code - agents plan, implement, and open pull requests from a specification - and keeps DevOps pipelines as its quality gates. The US Department of Defense uses "software factory" for DevSecOps teams and platforms such as Kessel Run and Platform One, which is a different meaning.

How do you build an AI software factory?

Start with a finished specification: a constitution, requirements with testable acceptance criteria, a solution design, and a dependency-mapped task list. Group related tasks into waves, run one isolated agent session per group, verify every pull request against its acceptance criteria and the constitution, merge one at a time under an agreed policy, write task status back to the spec, and stop at each milestone for human review.

Does a software factory replace software engineers?

No. The work moves up a level: engineers write and review specifications, decide trade-offs, design the verification, and approve milestones. Shapiro's level 3 puts it bluntly - you are not a senior developer any more, that is the AI's job; you are a manager.

Can AI-generated code ship without human review?

Some teams do it - StrongDM's rules are that code must not be written or reviewed by humans - but most teams keep a human gate. DORA's 2025 research found AI adoption still has a negative relationship with software delivery stability, and Addy Osmani warns of comprehension debt: the gap between how much code exists and how much any human understands. Checking each change against written acceptance criteria and reviewing at milestones keeps that debt small.

How do I run a software factory with Claude Code?

Install MySpec's Software Factory Manager plugin (myspec-factory) together with myspec-mcp. It turns one Claude Code session into a manager that plans waves from tasks.md, starts one Claude Code worker session per group of related tasks (in the cloud or in a git worktree), verifies and merges their pull requests, marks tasks done on MySpec, and stops at milestone gates for your review.

SECTION 10 SOURCES

Further reading

SPEC FILES FIELD GUIDE

The spec files a factory runs on

Every factory starts with drawings

MySpec's AI Architect interviews you and writes the spec bundle a software factory runs on: constitution.md, requirements.md with EARS+ acceptance criteria, solution.md, and a dependency-mapped tasks.md. Free during Open Beta.