Why measure software delivery?
Every organisation that builds or maintains software wants to know whether delivery is improving. Are changes reaching users faster? Are releases breaking things less often? When something does break, how quickly is it put right? Without shared measures, those questions get answered by impression, and impressions differ.
The most widely used measures come from DORA, a research programme run by Google Cloud that studies software delivery and its connection to organisational outcomes and individual well-being. Its research archive runs from 2014 to the 2025 report, which concluded that AI acts as an amplifier, with the greatest returns coming from attention to the underlying sociotechnical system.
From four keys to five metrics
Many people still refer to the “four keys”. The set has evolved, and DORA’s history of the metrics records the main changes:
- 2014: the first definition of IT performance used three metrics: deployment frequency, lead time for changes and mean time to recover (MTTR). Change fail rate was measured but did not correlate with the others, so it was left out.
- 2015: change fail rate joined the set, and the four were grouped into throughput (deployment frequency and lead time) and stability (MTTR and change fail rate).
- 2018 and 2021: DORA added availability as a measure of operational performance, later broadening it to reliability. DORA notes that its 2021 report wrongly described reliability as a “fifth metric”.
- 2023: MTTR was renamed failed deployment recovery time and redefined to cover only recovery from failures caused by a change to production, not external events such as a data-centre outage.
- 2024: deployment rework rate was added as a fifth delivery metric, to test whether change fail rate was acting as a proxy for rework.
So if you are setting up measurement today, the current model has five metrics, and reliability is measured separately.
The five metrics defined
DORA’s metrics guide groups the five into throughput and instability:
| Group | Metric | What it measures |
|---|---|---|
| Throughput | Change lead time | Time from a change being committed to version control to it being deployed in production |
| Throughput | Deployment frequency | Number of deployments in a period, or time between deployments |
| Throughput | Failed deployment recovery time | Time to recover from a deployment that fails and needs immediate intervention |
| Instability | Change fail rate | Proportion of deployments that need immediate intervention afterwards |
| Instability | Deployment rework rate | Proportion of deployments that are unplanned and happen because of an incident in production |
The two groups balance each other. Shipping more often while breaking more things is not improvement; neither is breaking nothing because almost nothing ships.
DORA also positions the metrics in time. It describes them as leading indicators of organisational performance and employee well-being, and as lagging indicators of software development and delivery practices. In other words, they tell you after the event whether your engineering practices are working, and they give early warning of effects on the wider organisation.
Speed and stability are not a trade-off
A central finding of the research is that speed and stability tend to move together. DORA’s guide states that the metrics are correlated for most teams, and that top performers do well across all five while low performers do poorly across them. The practices that make frequent, small changes possible, such as automation, testing and small batches, are the same practices that make each change safer.
This is the most useful message for anyone who assumes that more care must mean slower releases. Large, infrequent releases bundle many changes together, which makes them harder to test, harder to diagnose when they fail, and slower to fix.
How to measure the metrics in practice
Measure per application or service
DORA says the metrics are meant to be applied at the application or service level, one at a time, in the context of what the team is delivering. Averaging across an entire organisation blends very different systems and hides the signal.
Agree definitions first
Before collecting data, write down what counts for your service:
- Deployment: a release to production, including configuration changes if they carry risk.
- Commit time: usually the time the change was merged or committed to the main branch.
- Failure needing immediate intervention: a rollback, hotfix, feature-flag disable or emergency patch.
- Recovery: the point at which users are no longer affected.
- Rework deployment: an unplanned deployment made to address a production incident.
Consistent definitions matter more than precision. A metric that is calculated the same way every month shows trends clearly, even if the absolute numbers are imperfect.
Find the data
Most of the data already exists. Version control records commits and merges; CI/CD pipelines record deployments; incident and ticketing tools record failures and recovery. DORA’s guide suggests starting with conversations and its Quick Check, a five-question survey that takes under a minute and compares your answers with the rest of the industry, before investing in custom integrations. Commercial and source-available tools with pre-built integrations can automate collection later.
Google’s open-source Four Keys project, an earlier tool for collecting the metrics automatically, was archived in January 2024 and is no longer maintained. It also predates the current five-metric model, so treat it as a reference design rather than a ready-made solution.
Keep reliability in view
The delivery metrics tell you how well changes flow, not whether the service meets users’ expectations. For that, use service level objectives. Google’s Site Reliability Engineering book defines a service level indicator (SLI) as a carefully defined quantitative measure of some aspect of the service, such as latency or error rate, and a service level objective (SLO) as a target value or range for that indicator. It advises keeping as few SLOs as possible and not aiming for 100%. Instead, agree an error budget that sets how much unreliability is acceptable, and use the remaining budget to inform release decisions.
Together, the five delivery metrics and a small set of SLOs give a balanced view: how quickly and safely you change the system, and how well it serves users between changes.
Pitfalls to avoid
DORA’s guide is explicit about how the metrics go wrong:
- Turning metrics into targets. Mandates such as “deploy several times a day” encourage teams to game the numbers, an example of Goodhart’s law.
- Relying on a single metric. Use several, including some that pull against each other.
- Comparing very different applications. A regulated core banking system and a marketing website will rightly differ.
- Using the industry as an excuse. Regulation or sector norms are not a reason to stop improving.
- Siloed ownership. Giving separate teams separate metrics invites finger-pointing.
- Competing against other teams. The aim is your own improvement, not a league table.
- Measuring instead of improving. Expensive measurement infrastructure is not the goal.
How to improve the numbers
The metrics are outcomes. To move them, work on the capabilities underneath. DORA’s capability catalogue includes, among others:
- Version control and working in small batches
- Continuous integration and trunk-based development
- Deployment automation and continuous delivery
- Test automation and test data management
- Streamlining change approval
- Monitoring and observability, and proactive failure notification
- Loosely coupled teams and a generative organisational culture
DORA’s guide singles out reducing the batch size of changes as a common way to improve all five metrics at once. Smaller changes are quicker to review, test and deploy, less likely to fail, and easier to roll back when they do.
A simple way to start
- Pick one important service and gather its cross-functional team.
- Take the Quick Check to set a baseline.
- Agree definitions for deployment, failure, recovery and rework.
- Map how a change travels from commit to production and find the biggest constraint.
- Improve that one constraint, then measure again.
- Review the metrics in team retrospectives, alongside SLOs and user feedback.
Used this way, DORA metrics become a tool for conversation and steady improvement rather than a scorecard, which is how the research intends them to be used.