Engage

12 Aug 2026 · The software factory practice

A factory with the lights on

A software factory is the machine that turns what a business says into software it can use. We build one per customer, because the estate decides which agents are worth defining — and we run every one of them with the lights on, because the judgment is the part that cannot be automated.

What a software factory is

A software factory is an orchestrated set of AI agents that moves software work through stations: transcribing what was asked, drafting the specification, probing the systems it must touch, implementing the change, attaching the evidence that it works. Work moves through stations, every station leaves a record, and nothing leaves the building uninspected. The word factory is earned rather than borrowed.

The unit of work is the task. A task is opened from what the business actually said — a call, a ticket, a voice note — and it carries its origin with it, verbatim. It then moves through stages, agents writing what they did to the record at every one of them, people deciding against that record. One task, end to end, looks like this:

business userthe factorythe recordengineerthe ask — a call, tuesdaytask 214 opened · transcript attachedspecification drafted, for reviewspec approved · decision loggedloop — until the evidence holdsslice built · change 86 pushedevidence · 34 checks · screenshotdiff read · signed to shipin production · task 214 closedbusiness userthe factorythe recordengineer
Fig. 01 · One task, end to end. Agents (dashed rings) carry the work and write everything to the record; people (solid rings) decide against it. Red is the decision that ships.

As a technology services company, we do not sell that description. We stand the factory up inside the customer's account, tuned to their estate, and we run it with them — through the build and then for years of operation. The factory is part of what an engagement leaves behind: a machine for turning what the business says into software it can use.

One thing we have learned to say early: our factories are never lights-off. The industry has spent the past year trying to run factories with no one inside — agents planning, building, reviewing and merging on their own — and the published attempts keep converging on the same lesson. The volume holds up; the quality does not. A model can write more code every month, and it still cannot be trusted to judge whether a codebase is getting easier or harder to change, which is the judgment that decides whether the factory is compounding an asset or a liability. So the judge stays human, and the lights stay on.

One factory per customer

The second thing we say early is that we cannot supply the same factory twice. A factory is tuned to one estate's distribution of work — what arrives in a normal week, which systems it lands on, what evidence proves a change is safe there, and which changes are risky. That tuning is the factory. Remove it and what remains is a diagram.

Enterprise estates do not share a distribution of work. One organisation's week is feature requests against a portal; another's is integration breakage between a policy system and a warehouse, regulatory change with a filing date, and report requests that each take an analyst a day. The work arrives under different constraints — data residency, an air-gapped network, five legal entities answering to three regulators — and at different levels of maturity, from a team commissioning its first custom system to one that already reviews agent-authored pull requests every morning. The stations are stable; the agents are not.

How we implement one

Standing up a factory begins the way an architecture engagement begins: by reading. We take a recent quarter of the customer's actual work — requests, incidents, changes, escalations — and sort it. Which classes of work recur. Which are mechanical once the intent is clear, and which genuinely need judgment. Where the organisation is on its custom-software journey, and what its teams are ready to review.

Agents are then defined one class of work at a time, against a simple test: the work recurs, its output can be verified, and evidence of that verification can travel with the change. An insurer's factory opens with agents that transcribe intake and assemble compliance evidence. An organisation mid-migration gets agents that probe the old system's behaviour and carry fixes across versions. A company on its first custom build might start with only two: one that turns stated requirements into a reviewable specification, and one that attaches test evidence to every change.

Then, before the volume is switched on, the gates go in. Every factory we run has four places where a person decides: what to make and why, before anything is funded; how it fits the enterprise architecture, before anything is designed; the shape of the code — the types, the interfaces, the call paths — before anything is implemented at scale; and the review that signs a change to ship. Agents carry the volume between the gates. Nothing passes a gate without a person.

Because every stage must leave something a person can open, the factory's output is not a stream of diffs — it is a chain of artifacts, each one made to be judged in minutes:

task 214the ask verbatim+ transcriptspecificationscope · the planone open questionchange 86diff · 34 checksscreenshotthe decisionrs · 12 augsigned to shipevery stage leaves something a person can open
Fig. 02 · Every stage leaves an artifact: the ask verbatim, the reviewable specification, the change with its evidence, the recorded decision. The rings say who made each.

The factory also builds in slices rather than in layers. A slice is a thin, working path through the whole system — something a business user can touch in a browser while the build is days old, not months. Slices exist because of an old truth that agents have made newly expensive to ignore: the cost of changing your mind rises with every step away from the plan. A decision argued over a one-page design costs minutes. The same decision discovered in review, under two thousand lines of finished code, costs days — and in production it costs an incident.

How we keep the lights on

Keeping the lights on does not mean watching dashboards. It means the judgment never switches off, because a factory left to judge itself degrades quietly: each change works, each change also makes the next one a little harder, and no single week is the week it went wrong. Four disciplines prevent that.

A person reads the code. Not a model grading a model — an engineer scrolling the diff, holding the codebase's hard-won opinions about how software should be. Planning happens in documents that are argued before they are built, so review is confirmation rather than archaeology. Every change carries its evidence with it, so the reviewer judges the change instead of reconstructing it. And the record — the living documentation, the decisions, the transcripts — is kept current, because it is what the factory's agents read; a factory running on a stale map of the estate produces confident work in the wrong place.

The factory itself is also operated, not just used. Agents are versioned and re-tuned as the estate changes, their output is audited against the record, and the weekly steering call reads outcomes — what moved, what did not — so the factory's next week is steered by results rather than momentum.

The spine of all four disciplines is the decision log. Every ruling a person makes — at a gate, on a call, in a review — is recorded with its date, its author and its reason, and the agents read it back before they build. A reason outlives the meeting it was said in:

decisions · broker portal programme05 augfactory opens with intake + evidence agents08 augspec narrowed to the bind flow — renewals wait11 augrating service untouched: vendor freeze to oct12 augchange 86 signed to shipthe agents read this before they build the next thing
Fig. 03 · The decision log. Dated rulings with a reason and a name — the file the agents read before they build, and the file that answers 'why is it like this?' in year three.

What a human does, and what we do not do

The honest division of labour is now stable enough to write down. People decide what to make and why, and stand behind it to the business. People own the architecture and the program design — the decisions that are cheap on paper and ruinous in production. People read the code, take the steering call, and sign what ships; a Damco engineer's name is on every release. And people hold the relationship: the strategist who was in the discovery call is the one accountable in week twelve.

What people no longer do is the grunt work. Re-typing what a document already says, scaffolding, the fourth integration probe of the month, assembling regression evidence, carrying a fix across four versions, chasing status that the record can answer — that is the factory's volume, and every hour it absorbs is an hour an engineer spends on judgment instead.

The reason a person can hold four gates across a factory's whole throughput is that the factory prepares every decision. Nothing arrives as two thousand raw lines; it arrives as an artifact built to be judged — what changed, the evidence, and the one question that needs a human answer:

change 86 · ready for reviewtask 214 · broker portal · prepared by the build agentwhat changedthree files · the bind flow onlyevidence34 checks · screenshot · staging runthe questionbrokers bind without re-keying. ship?signed to ship · rs · 12 augdecision logged to the record, with its reason
Fig. 04 · The review, prepared. The factory assembles what changed, the evidence, and the one question — a decision a person makes in minutes, recorded forever.

There is also a list of things the factory is not allowed to do, and it is short and absolute. It does not decide what to build. It does not merge. It does not judge its own quality. And it does not ship a change on the strength of its own tests alone — evidence is attached for a person to weigh, never accepted on the factory's word.

Why this is worth it

The factory exists for two movements every CTO is already trying to make. Business users get to software sooner: the mechanical majority of a change stops queueing behind engineering capacity, and a stated need becomes a reviewable, evidenced change in days. And engineers stop being spent on volume: their hours move to the four gates, which is where quality is actually decided.

There is a third movement, quieter than the first two. We have written before about the compression problem: software gets built from a description of a description, and what falls out at each retelling is usually what mattered. The factory is the mechanical answer. The business user's words are captured once — the call transcribed, the transcript attached to the task it produced — and the agents that build from that task read the original, not a paraphrase of it. What was said becomes a specification, the specification becomes code, and the record travels with the change into production, where the same factory reads it again when the system needs to evolve.

See what a build week looks like →