When a leadership team tells us delivery feels slow, they usually cannot say where the time goes. Story points are up, sprints complete, and somehow the roadmap has slipped a quarter. The measures being tracked describe activity rather than flow.
Four numbers, drawn from a decade of delivery research, describe the thing leaders actually care about: how quickly work reaches users, and how safely.
The four
- Lead time for change: from the first commit to running in production. Not the estimate; the elapsed clock.
- Deployment frequency: how often you ship. Weekly is fine. Quarterly is a diagnosis.
- Change failure rate: the share of releases that need a fix, rollback, or hotfix.
- Time to restore: how long recovery takes when something does break.
What makes these useful is that they resist gaming. You cannot improve lead time by working longer hours, and you cannot improve change failure rate by shipping less carefully. Both require fixing the system: environments, review queues, test reliability, and the size of the batches being released.
Where the time actually goes
In every delivery diagnostic we have run, the waiting dominates the working. Code sits in review for three days because the only person who can approve it is in workshops. A release waits eleven days for a manual regression pass. A branch cannot be tested because staging is occupied by another team's migration.
Nobody is idle and nothing is moving. That is the signature of a queue problem, not an effort problem.
Measure the elapsed time in each state and the bottleneck usually announces itself within a week. It is almost never that engineers are typing too slowly.
Smaller batches fix more than they should
The most reliable intervention we know is reducing the size of what gets released. Smaller changes are easier to review, safer to deploy, faster to diagnose, and cheaper to revert. Improve batch size and all four numbers tend to move together, which is a good sign you are treating a cause rather than a symptom.
One caution
Do not turn these into individual performance measures. The moment lead time appears on a personal scorecard it stops describing the system and starts describing what people are willing to report. Keep them at team level, review them monthly, and use them to argue for investment in the plumbing (the environments, the test suite, the pipeline) that never quite wins a prioritisation meeting on its own merits.
About the author
Angel Maile
Angel co-founded Bonang Technologies after a decade spent building software inside organisations where the technology decisions and the commercial ones were made in separate rooms. The company exists to close that gap: engineering that starts from what the business is actually trying to achieve, and leadership that can hold both conversations at once.