Measuring software team productivity with agile metrics

No single agile metric measures the productivity of a software development team, and the ones most often used for the purpose, velocity and burndown, were designed for planning. The metrics that tell you most are the ones that follow a piece of work from request to live use: how long it takes, how many items are finished, and how often a release causes a problem. Used as trends for one team over time they are informative. Used as targets, or to compare teams or individuals, they stop being true.

That last point is the reason productivity is hard to measure. Lines of code, hours logged and tickets closed all count activity, and each can be increased without delivering anything of value. Now that AI coding assistants can produce large amounts of code in minutes, counting output is weaker evidence than it has ever been. What a business wants is working software in use, delivered at a predictable pace, that does not break.

Planning metrics: velocity and burndown

Many teams estimate each piece of work in story points, a relative measure of size. An item given 5 points is judged to be roughly as much work as other 5-point items the team has done. Velocity is the number of points the team completes in a Sprint (a fixed working period of a month or less).

Velocity is useful to the team itself. If it has finished between 30 and 40 points in each of the last six Sprints, it should not plan 60 for the next one, and a backlog of 300 points is about eight to ten Sprints of work.

It is not a measure of productivity:

  • Points are the team’s own judgement, so two teams’ velocities cannot be compared.
  • If management asks for velocity to rise, estimates grow and the number rises with no change in what is delivered.
  • It counts effort completed, not value delivered.

Neither story points nor velocity appears in the Scrum Guide (2020). They are optional practices.

A burndown chart plots the work remaining in a Sprint or a release against time. The Scrum Guide mentions burn-downs only as one of several practices for forecasting progress. A burndown is a daily early warning for the team that the Sprint Goal is at risk, and it says little once the Sprint is over.

Flow metrics: cycle time and throughput

Flow metrics come from Kanban, and they are the most useful measures for a customer because they are expressed in days and counts. The Kanban Guide (May 2025 edition) defines four:

  • Cycle time: the elapsed time between when a work item started and when it finished.
  • Throughput: the number of work items finished per unit of time.
  • Work in progress (WIP): the number of items started but not finished.
  • Work item age: the elapsed time between when an item started and today.

Cycle time answers the question a business asks most: if we request a change, how long until we have it? Look at the spread as well as the average. A team whose changes usually take four days but sometimes take forty has a predictability problem that an average hides.

Rising work in progress with flat throughput means the team is starting more than it finishes. Work item age picks out the items that are stuck, while there is still time to do something about them. Our article on reasons to use Kanban explains the method behind these measures.

Delivery metrics from DORA

DORA is a research programme, now run by Google Cloud, that has studied software delivery since 2014. Its guidance currently defines five software delivery metrics, in two groups.

Throughput:

  • Change lead time: how long a change takes to go from being committed to version control to being deployed in production.
  • Deployment frequency: how often changes are deployed.
  • Failed deployment recovery time: how long it takes to recover from a deployment that fails and needs immediate intervention.

Instability:

  • Change fail rate: the proportion of deployments that need immediate intervention afterwards, such as a rollback or an urgent fix.
  • Deployment rework rate: the proportion of deployments that were unplanned and happened because of an incident in production.

If you learned these as “the four key metrics”, the list has changed. Mean time to recover was renamed and redefined as failed deployment recovery time in 2023, and rework rate was added in 2024.

DORA’s own guidance carries three warnings that are worth repeating. The metrics are meant to be applied to one application or service, and comparing very different applications is misleading. Setting them as targets increases the likelihood that teams will game them. And speed and stability are not a trade-off: in its research, teams that release more often tend to be more stable too. These metrics depend on the automated build and release practices described in what is DevOps.

What each metric can and cannot tell you

MetricGood forNot for
VelocityThe team’s own planning and forecastingComparing teams, or judging output
BurndownSpotting a Sprint at risk while it is runningAnything after the Sprint ends
Cycle timePredicting how long a request will takeJudging an individual developer
ThroughputSeeing whether the pace is steadyMeasuring value, since items differ in size
Change fail rateSeeing whether releases are safeA target to be hit by releasing less often
Defects found after releaseJudging quality as users experience itBlame, since many defects start as unclear requirements

What to measure if you are the customer

If you employ the team, choose a small set that covers speed, stability and quality together, so that improving one at the expense of another shows up. A workable set is cycle time, throughput, change fail rate and the number of defects reported after release. Review the trend monthly with the team, and ask what is behind a change before drawing a conclusion from it.

Do not measure individuals with any of these. Software is built by teams, and the developer who spends a day unblocking three colleagues has a poor personal count and a large effect.

If you buy from a supplier, the team’s internal productivity is mostly the supplier’s concern, particularly on fixed-price work. What you need to see is different:

  • how many items were delivered and accepted against their acceptance criteria in the period;
  • how long change requests take from approval to live use;
  • how many defects were found after release, and how quickly they were fixed;
  • for a support agreement, response and resolution times against what was agreed. Our page on software maintenance explains how support plans are structured.

None of these needs special tooling. The tracking tools that agile teams already use, covered in top tools for agile teams, record when each item started and finished, and that is enough to produce every flow metric above.

A first step

Ask your team or supplier for one figure: for the last twenty changes delivered, how many days passed between the request being approved and the change being live. If they can produce it within a day, the basics of measurement are in place. If they cannot, that is the first thing to fix, and it matters more than which metric you pick. When you are comparing suppliers, our questions to ask a software supplier include one on how you will see progress.

Tell us about your system

Say what it does, what it is built on and what is worrying you. We will reply with what we would look at first and whether we are the right people to help.

Tell us about your system 0800 433 7990 Monday to Friday, 9am to 5pm. A first 20-minute call is free, and we reply to every enquiry within one working day. What happens after you get in touch