Your engineers got faster.
Your delivery didn’t.
We ran a full diagnostic against two years and 5,316 work items of a real 7-team engineering department. Output more than doubled. The slowest work got three times slower. Nobody noticed, because every dashboard said things were improving.
Analysis · 7 product teams · 5,316 items · 24 months · metadata only
Every engineering leader we talk to has bought the tooling. Copilot, better pipelines, more automation. And most of them can point at a chart showing output going up.
Then the roadmap slips anyway, and nobody can explain it.
We wanted to know where the gain goes. So we took a real payroll technology engineering department, seven product teams, two years of history, every ticket, and measured it end to end. Not a survey. Not a maturity questionnaire. The actual timestamps.
What we found
The average item took 31.9 working days from start to finish.
Only 5.7 of those days was anyone touching it.
Where a typical work item spends its life
Meanwhile the teams were getting better. Completed items per quarter went from 291 to 686 over the period, a 136% increase. Over the same stretch, their slowest work went from 25 working days to 65, and at one point 135.
One definition, then we will keep it plain. When we say “the slowest work” we mean the point where 85 out of 100 items are finished. It is the number worth watching, because the last 15 are the ones that blow up a plan. The typical item is not what makes you miss a date.
That is the trap. Averages hide the slow work, and commitments are made against the slow work. Nobody plans around their typical ticket. They plan around the feature that has to ship in March.
The bottleneck wasn’t where anyone thought
Going in, the assumed constraint was QA. It usually is. The data said otherwise.
| Stage | Days per item | What it is |
|---|---|---|
| Ready for Release | 6.4 | Waiting |
| Ready for QA | 5.3 | Waiting |
| To Do | 5.2 | Waiting |
| Estimate | 3.9 | Waiting |
| Ready for Code Review | 3.5 | Waiting |
| In Development | 2.6 | Being worked |
| In QA | 2.3 | Being worked |
| Code Review | 0.6 | Being worked |
Read the two QA rows together. Work waits 5.3 days to be picked up, then clears QA in 2.3. It is tempting to call that a queue problem rather than a QA problem, but QA owns that queue. So the honest reading is 7.6 days of QA-owned time per item, most of it before anyone starts. The same shape sits in front of code review: 3.5 days waiting, 0.6 days reviewing.
Which reframes the whole thing. Development is producing faster than the rest of the system can absorb, and the queues are where that mismatch piles up. The answer is not to tell anyone to work harder. It is to balance the line: hold work in progress to what the next stage can take, raise that stage’s capacity, or both. That ratio is set above the team.
And the single largest pool of wasted time was the one nobody would have guessed. Finished, tested, accepted work sitting in “Ready for Release” for 6.4 days. A fifth of all the time in the system. That is not a capacity problem. It is a calendar problem.
Two more things the data gave up
Half the defect load had a named upstream source
824 of 1,671 defects were linked to a support ticket. Those defects took 62% longer to close than the ones raised internally: 42 working days against 26, measured at the slow end of the range. On the worst performing team it was 68% of their defect load and their slow work ran to 95 days.
Support had supplied 40 to 51% of all defects every single quarter for two years. It was not a spike anybody would notice. It was the baseline, and nothing filtered it before it hit a team board.
A feature was a container, not a commitment
Across 164 measurable features, the typical one ended up 1.5 times the size it was two weeks in. The worst quarter grew 4 times or more. Three in four had work added after the first piece had already shipped.
The easy read is that stakeholders kept piling work on. The data does not support that. This org runs Kanban, and nothing in Kanban marks the point where a feature stops taking work. Scope kept arriving because no boundary told it to stop. That is a process gap, not a discipline problem, and it is worth knowing which one you have before you try to fix it.
We tested what that does to a forecast. One real feature in flight, 24 items remaining, team finishing about 6 a week:
| Forecast method | 85% confident by |
|---|---|
| Scope as it appears on the board | 6 weeks |
| Scope adjusted for how features actually grow | 18 weeks |
Twelve weeks of difference, and all of it is work that is not on the board yet. This is why estimates keep failing in organizations where nobody estimates badly.
What we’d change, and what it’s worth
| Change | Effort | Days saved per item |
|---|---|---|
| Release more frequently | Low. A calendar decision. | 5.4 |
| Cap the work waiting at each handoff | Medium. Daily discipline. | 6.9 |
| Put a triage gate on support intake | Medium. Needs a named owner. | 3.5 |
| Forecast using how features actually grow | Low. Changes the question asked. | fewer misses |
Roughly 32 working days down to about 14. Same people. The gain comes from removing waiting, not from adding capacity.
Why we’re publishing the limits too. This analysis still cannot explain why support work is slower. It clears the release queue faster than internal work, so the delay is somewhere we have not looked yet. A ticket also loses its owner the moment it enters a waiting column, so the board can say who is working but not who owns anything stuck. And three of four shared services report completion so unreliably that a first pass produced a confident, specific, wrong finding, one we caught only by cross checking against a second field. A diagnostic that never says “I don’t know” is not a diagnostic. It is a sales document.
If any of this sounds familiar
The pattern repeats. Tooling investment lands, output improves, delivery does not, and the constraint turns out to be a queue nobody was measuring, because no dashboard shows how long things have been sitting still.
You can check your own numbers in about five minutes. The question worth asking first is simple. What share of your delivery time is spent waiting? If you cannot answer that from your current reporting, that is the finding.
Find your own 82%
A free Delivery Health Check gives you the first read. No export required.
Start the Health Check →Method. Read only, metadata only extraction: dates, types, statuses and links. No ticket text was read or stored at any point. Times are in working days, abandoned items excluded. Where we describe “the slowest work” we mean the point where 85 out of 100 items are complete, used throughout because the spread is wide on every team. Stage timing came from status change history on a 140 item sample of the slowest defects, split evenly between support originated and internally raised. Completion data was validated against a second status field before use. The forecast was built from 164 features with measurable scope history and 78 weeks of delivery rate, sampled 20,000 times.
This is an analysis of a real engineering organization, published with its identity withheld. It is not a client case study.