Findings · Delivery Intelligence

Your engineers got faster.
Your delivery didn’t.

We ran a full diagnostic against two years and 5,316 work items of a real 7-team engineering department. Output more than doubled. The slowest work got three times slower. Nobody noticed, because every dashboard said things were improving.

Analysis · 7 product teams · 5,316 items · 24 months · metadata only

Every engineering leader we talk to has bought the tooling. Copilot, better pipelines, more automation. And most of them can point at a chart showing output going up.

Then the roadmap slips anyway, and nobody can explain it.

We wanted to know where the gain goes. So we took a real payroll technology engineering department, seven product teams, two years of history, every ticket, and measured it end to end. Not a survey. Not a maturity questionnaire. The actual timestamps.

What we found

82%
of a work item’s life is spent waiting in a queue.
The average item took 31.9 working days from start to finish.
Only 5.7 of those days was anyone touching it.
82% WAITING
18% WORK

Where a typical work item spends its life

Meanwhile the teams were getting better. Completed items per quarter went from 291 to 686 over the period, a 136% increase. Over the same stretch, their slowest work went from 25 working days to 65, and at one point 135.

One definition, then we will keep it plain. When we say “the slowest work” we mean the point where 85 out of 100 items are finished. It is the number worth watching, because the last 15 are the ones that blow up a plan. The typical item is not what makes you miss a date.

Output up 136%. The slowest work up 159%. And the middle of the range barely moved, 14 days to 18, so every report built on averages showed a department steadily improving.

That is the trap. Averages hide the slow work, and commitments are made against the slow work. Nobody plans around their typical ticket. They plan around the feature that has to ship in March.

The bottleneck wasn’t where anyone thought

Going in, the assumed constraint was QA. It usually is. The data said otherwise.

StageDays per itemWhat it is
Ready for Release6.4Waiting
Ready for QA5.3Waiting
To Do5.2Waiting
Estimate3.9Waiting
Ready for Code Review3.5Waiting
In Development2.6Being worked
In QA2.3Being worked
Code Review0.6Being worked

Read the two QA rows together. Work waits 5.3 days to be picked up, then clears QA in 2.3. It is tempting to call that a queue problem rather than a QA problem, but QA owns that queue. So the honest reading is 7.6 days of QA-owned time per item, most of it before anyone starts. The same shape sits in front of code review: 3.5 days waiting, 0.6 days reviewing.

A queue that never drains is not waste. It is the signature of a stage receiving more than it can process.

Which reframes the whole thing. Development is producing faster than the rest of the system can absorb, and the queues are where that mismatch piles up. The answer is not to tell anyone to work harder. It is to balance the line: hold work in progress to what the next stage can take, raise that stage’s capacity, or both. That ratio is set above the team.

And the single largest pool of wasted time was the one nobody would have guessed. Finished, tested, accepted work sitting in “Ready for Release” for 6.4 days. A fifth of all the time in the system. That is not a capacity problem. It is a calendar problem.

Two more things the data gave up

Half the defect load had a named upstream source

824 of 1,671 defects were linked to a support ticket. Those defects took 62% longer to close than the ones raised internally: 42 working days against 26, measured at the slow end of the range. On the worst performing team it was 68% of their defect load and their slow work ran to 95 days.

Support had supplied 40 to 51% of all defects every single quarter for two years. It was not a spike anybody would notice. It was the baseline, and nothing filtered it before it hit a team board.

A feature was a container, not a commitment

Across 164 measurable features, the typical one ended up 1.5 times the size it was two weeks in. The worst quarter grew 4 times or more. Three in four had work added after the first piece had already shipped.

The easy read is that stakeholders kept piling work on. The data does not support that. This org runs Kanban, and nothing in Kanban marks the point where a feature stops taking work. Scope kept arriving because no boundary told it to stop. That is a process gap, not a discipline problem, and it is worth knowing which one you have before you try to fix it.

We tested what that does to a forecast. One real feature in flight, 24 items remaining, team finishing about 6 a week:

Forecast method85% confident by
Scope as it appears on the board6 weeks
Scope adjusted for how features actually grow18 weeks

Twelve weeks of difference, and all of it is work that is not on the board yet. This is why estimates keep failing in organizations where nobody estimates badly.

What we’d change, and what it’s worth

ChangeEffortDays saved per item
Release more frequentlyLow. A calendar decision.5.4
Cap the work waiting at each handoffMedium. Daily discipline.6.9
Put a triage gate on support intakeMedium. Needs a named owner.3.5
Forecast using how features actually growLow. Changes the question asked.fewer misses

Roughly 32 working days down to about 14. Same people. The gain comes from removing waiting, not from adding capacity.

Why we’re publishing the limits too. This analysis still cannot explain why support work is slower. It clears the release queue faster than internal work, so the delay is somewhere we have not looked yet. A ticket also loses its owner the moment it enters a waiting column, so the board can say who is working but not who owns anything stuck. And three of four shared services report completion so unreliably that a first pass produced a confident, specific, wrong finding, one we caught only by cross checking against a second field. A diagnostic that never says “I don’t know” is not a diagnostic. It is a sales document.

If any of this sounds familiar

The pattern repeats. Tooling investment lands, output improves, delivery does not, and the constraint turns out to be a queue nobody was measuring, because no dashboard shows how long things have been sitting still.

You can check your own numbers in about five minutes. The question worth asking first is simple. What share of your delivery time is spent waiting? If you cannot answer that from your current reporting, that is the finding.

Find your own 82%

A free Delivery Health Check gives you the first read. No export required.

Start the Health Check →

Method. Read only, metadata only extraction: dates, types, statuses and links. No ticket text was read or stored at any point. Times are in working days, abandoned items excluded. Where we describe “the slowest work” we mean the point where 85 out of 100 items are complete, used throughout because the spread is wide on every team. Stage timing came from status change history on a 140 item sample of the slowest defects, split evenly between support originated and internally raised. Completion data was validated against a second status field before use. The forecast was built from 164 features with measurable scope history and 78 weeks of delivery rate, sampled 20,000 times.

This is an analysis of a real engineering organization, published with its identity withheld. It is not a client case study.

← All findings