Press ESC to close

How Do You Monitor an AI Pipeline That Runs Without You?

Plenty of engineering teams now run pipelines where AI agents pick up a ticket, write the code, review it, test it and open the pull request, with a person involved at a few points and not at the rest. Dan Shapiro gave the far end of that a name in January: the dark factory, borrowed from manufacturing plants that run without lighting because nothing inside them needs to see. StrongDM published theirs a fortnight later. A good many teams are further along than they say in public.

Suppose you’ve built one and it’s been running since the spring with nothing going wrong. Ask what evidence you have that it’s still doing what you designed it to do, and the honest answer is the absence of bad news. You’d observe the same quiet either way. You believe you’re governed, you report that you’re governed, and you have no evidence in either direction.

The limits you set on a line like that are claims about behaviour. This agent writes to the source folder and not the test folder. That one stops and asks before it spends money. A third gets a fixed budget of turns, meaning individual goes at the problem, and stops when it runs out whether or not it has finished. Call them walls, because that’s what they are: the boundary round the whole line, and the boundaries between the stations inside it.

Claims decay. What tells you whether yours are holding is what the line tried to do and was stopped from doing.

The wall that moved

Nobody moves a wall. It gets moved one reasonable permission at a time.

A release is stuck on a Friday because the agent writing the code can’t create the test fixture it needs, so somebody widens what it’s allowed to write to. That’s the right call on the day. It’s also the call that put an agent’s hands on the files that check its own work, and the design document saying that agent writes to the source folder and not the test folder was never opened, because nobody felt like they were changing the architecture. They were unblocking a release.

Do that a few times in a quarter and the walls are somewhere new, with no record of the move. Each step was a small operational decision, and the boundary was written down somewhere that small operational decisions don’t reach.

You’ll not find this by reading the design document. The document still describes the walls you drew. That’s what documents do.

Your dashboards stopped measuring what they used to measure

DORA has been tracking the industry-level version of this. Their 2024 research found that every 25% rise in AI adoption came with an estimated 7.2% drop in delivery stability, meaning more deployments that fail and take time to put right, and a 1.5% drop in throughput, the amount of change reaching production at all. By 2025, with 90% of nearly five thousand respondents using AI at work, throughput had turned positive. Stability had not.

Their 2026 ROI report names two of the reasons. The verification tax is the time developers spend checking generated code, which grows with the volume of it. Pipeline adaptation is what happens next, as testing and change approval have to scale to absorb work arriving faster than they were built for. Generating code got cheap. The stages that check it did not, so that’s where the work piles up.

Put that next to the stability number and the problem’s in plain sight. Deployment frequency was a good number because volume tracked effort, and effort tracked a person deciding something was ready to go. Generation broke that coupling. The number still climbs and it’s stopped meaning what it meant. DORA got to the same place from another direction, describing AI as an amplifier that magnifies whatever an organisation already had. Your line doesn’t acquire new properties. It runs the ones you built, faster, and past the point where a person was absorbing the difference.

Measuring the wall, not the traffic

A dashboard of agent runs tells you the line is busy. Busy was never the question.

Whether the line stayed inside its walls is a different measurement about a different thing. Traffic metrics are about the work: how much, how fast, how often. Boundary metrics are about the wall: what it refused, how close anything got, how often it was widened and who widened it.

Four signals carry most of it, and a line built correctly produces all four already, without anyone having to build something special to collect them: how much of its turn budget each station uses; how often a station was stopped from reaching somewhere it shouldn’t go, and where it was heading; how often a review gate sends work back; how often work gets escalated, and to whom.

None of those is a productivity measure and none belongs on the same screen as one. Put them next to deployment frequency and somebody will read them as friction, which is the reading that gets them optimised away.

Limits before the run

A measurement with no limit attached is a number you’ll rationalise.

Refusals at the code-writing agent have doubled this month. Is that a wall doing its job against a model that’s got more ambitious, or a station being handed work it was never scoped for? You can argue either after the fact, and whichever you argue will be the one that fits the week you’re having.

Set the limit first. Manufacturing solved this a century ago with statistical process control, which is the practice of writing down what normal looks like before the run, then treating anything outside those limits as a thing to investigate rather than a thing to explain away. Article 9 makes the case for carrying it into software. The discipline is writing the number down before you have a reason to want it different.

For an autonomous line the limits worth setting in advance are the boring ones: what a station’s turn use looks like on ordinary work, how many refusals a week is normal for each wall, what share of changes a review gate sends back when it’s working properly. Take a fortnight of ordinary running, write down what you see, and you’ve got something to be surprised by.

The signal you’re throwing away

An agent reaching for a file it’s not allowed to touch is the most informative event your pipeline can produce. Most pipelines log it as a failure and retry.

That reflex comes from the era when a blocked action meant a misconfiguration, because the thing being blocked was a script doing what it was told. An agent selects its own actions while it runs. When it reaches somewhere it cannot go, the reach tells you what the model concluded the task required. The wall did its job, and the attempt is the finding.

Read a quarter of refusals and you learn things no amount of output review gives you. Which stations are being asked to do work outside their scope. Which walls are load-bearing and which have never once been tested. Where the model’s idea of the task and yours have quietly diverged.

One thing has to be true first. Only an enforced wall can refuse anything. Take the write tools away from a reviewer in its settings and it will refuse; tell it in its instructions not to write and one day it’ll write anyway, with nothing logged, because nothing was blocked. Your refusal count is a census of which walls exist.

Which is also where reading stops being enough. Going through a pipeline’s code, settings and prompts will tell you which walls are declared and which of those are enforced. Whether a stop halts, and how far a failure travels before something catches it, cannot be settled that way at all. An honest audit says so and writes a test instead.

Reading tells you a wall is declared. A refusal is the wall being exercised, and a wall that has never produced one is a claim nobody has tested.

What no record reconstructs

Every station can hold a perfect record of what it did and you’ll still not reconstruct what the line decided. No single station made that decision. It came out of how the planner framed the work, how the next agent read that framing, what the reviewer accepted and what the gate let through. No station held it, because no station could see it. Each one held its own part and did its part correctly.

Your AI Knows What It Did. Not What It Decided. described a single system that cannot tell you why it chose what it chose, where the reasoning exists and goes uncaptured. A line is the harder case. The decision no participant ever held was never anywhere to capture. The first is a problem with recording. The second is not, and a better logger will not touch it.

What does help is recording where the walls were, alongside the work. Which walls were in force for this run, which exceptions were granted and why, which refusals fired, what the limits were on the day. All of that can be reconstructed. The reasoning can’t, and pretending otherwise is how a governance programme ends up with archives nobody can answer a question from.

You drew the walls, you built them, and the line runs. What you have after a quiet quarter is a claim about behaviour and no evidence either way.

Walls were never a document. They were a claim, and claims need evidence.

Which leaves the thing none of this inspects. You can place every gate correctly, wall every station and measure every boundary, and the walls hold, the gates are where they should be, and every one of them is checking the work against something nobody has looked at.

Sources

  • Accelerate State of DevOps Report 2024, DORA, 2024. Source for the estimated 1.5% reduction in delivery throughput and 7.2% reduction in delivery stability for every 25% increase in AI adoption.
  • State of AI-assisted Software Development, DORA, 2025. Source for throughput turning positive while stability did not, for 90% of nearly five thousand respondents using AI at work, and for AI as an amplifier that magnifies the strengths of high performing organisations and the dysfunctions of struggling ones. Survey conducted 13 June to 21 July 2025.
  • ROI of AI-assisted Software Development, DORA, v2026.1, 11 May 2026. Source for the verification tax and pipeline adaptation, two of the three named causes of the J-Curve productivity dip. Worth knowing before you quote it: the sample calculator in its appendix is headed “for demonstration purposes only”, so the change failure rate moving from 5% to 6%, and the $344,000 of downtime that follows, are illustration values chosen to show the arithmetic. Both circulate widely as findings. They are not.
  • The Five Levels: from Spicy Autocomplete to the Dark Factory, Dan Shapiro, 23 January 2026. Coins the dark factory usage in software, taken from FANUC. Level 5 is the one this piece is about.
  • StrongDM Software Factory, Justin McCarthy, Jay Taylor and Navan Chauhan, StrongDM, February 2026. The first public account of a line running without human code review, and the reason the opening can say teams are further along than they admit in public.
  • Economic Control of Quality of Manufactured Product, Walter A. Shewhart, Van Nostrand, 1931. Source for setting limits before the run rather than arguing about them afterwards. The first control chart appeared in a one-page memorandum Shewhart wrote for George Edwards at Bell Telephone Laboratories in May 1924, which is where the century is counted from.

Related reading (Synaptic Pixels)

Leave a Reply

Your email address will not be published. Required fields are marked *