Diagnose the delivery system before you try to transform it

Why this practice diagnoses before it prescribes, what a delivery diagnostic can and cannot see, and the premises we are prepared to be wrong about.

Beyond Delivery Partners2 min readOrganizational change

Most engineering organizations that feel slow have already tried something. A framework, a reorg, a tooling migration, sometimes all three in eighteen months. The pattern that brought this practice into existence is what happens next: the feeling persists, the theories multiply, and the loudest theory wins the next budget cycle. Nobody lacks conviction. What is missing is a diagnosis.

So the first engagement here is always the same: four weeks reading how delivery actually happens in your organization, from your own pipeline data and the people who live in it. Before we prescribe anything, it seems fair to write down the premises we work from, so you can hold us to them.

The first premise: delivery problems are system problems. When a team slows down, the cause is almost never that the people got worse. The codebase aged, the review queue lengthened, the on-call rotation thinned, the release process accreted a checkpoint nobody remembers approving. Systems produce the outcomes; people cope with the systems. A diagnostic that ends by naming a team has failed.

The second: where time goes is an empirical question. Organizations argue about velocity in the abstract while their own tooling holds the answer in specific. Ticket timestamps, pipeline runs, review latencies, incident timelines. The data is imperfect and partial, which is why we pair it with interviews, but between the two the question "where do our weeks go" stops being a matter of opinion. In most organizations we have seen, nobody has ever actually looked.

The third: measurement is for the system, never for the people. The moment a delivery metric touches a performance review it stops measuring the system and starts measuring fear. Our engagement letters rule it out in writing, not as a courtesy but because the diagnostic depends on people telling us the truth, and people do not tell the truth into an instrument that can rank them.

The fourth: most findings are boring, and that is good news. The public conversation about engineering effectiveness runs to frameworks and philosophies. The queues we actually find are concrete: a review step with one qualified approver, a test suite that takes forty minutes and fails one run in five for reasons nobody investigates, an environment that three teams share and each believes the others own. Boring findings have executable fixes. Interesting findings mostly have book deals.

The last premise is about us. A diagnosis you cannot execute without the diagnostician is a subscription, not a finding. Everything the review produces is written to work after we leave: evidence attached, first moves attached, no dependency on our calendar. Some clients ask us to stay for the first move. The design goal is that none of them have to.

That is the practice. The lessons that follow in this space will be the patterns we keep meeting, published with the evidence, because the reasoning is the product and you should get to inspect it before you buy any.