The importance of code reviews, and what they catch

A code review is a second developer reading a change before it is accepted into the system. Code reviews matter for two reasons. They catch mistakes the author cannot see in their own work. They also ensure that more than one person understands each part of the system, and for a business that depends on its software, the second reason is often worth more than the first.

What a code review looks like now

A developer finishes a change and submits it as a pull request, which is a proposal to merge the change into the main copy of the code. This happens in the tool where the code is stored, such as GitHub or Azure DevOps. Automated checks run first to confirm that the code builds and the tests pass. A colleague then reads the change line by line, asks questions, requests corrections and approves it. Both tools can be set to block a change until a reviewer has approved it, and the whole exchange is kept as a record against the change.

This is much lighter than the formal code inspections of earlier decades, which involved meetings and printed listings. Many teams now add an automated review, including AI tools that comment on a pull request, as a first pass. These are useful for slips and inconsistencies. They do not know your business, and a tool cannot be the second person who understands the system.

What the research shows

Three sources are worth knowing, along with their limits.

Barry Boehm and Victor Basili, writing in the journal IEEE Computer in January 2001, summarised the studies then available and reported that peer reviews caught between 31 and 93 per cent of defects, with a median of around 60 per cent. Those studies predate the lighter, tool-based style of review used now. The same article observes that reviews, analysis tools and testing “catch different classes of defects”, so none of them replaces the others.

Alberto Bacchelli and Christian Bird studied modern, tool-based review at Microsoft and published the results at the International Conference on Software Engineering in 2013. They observed and surveyed developers and managers and classified hundreds of review comments. Their finding was that although finding defects is the main reason teams give for reviewing, reviews “are less about defects than expected” and instead bring knowledge transfer, wider awareness within the team and alternative solutions to problems.

SmartBear, which sells review software, published a study of a programming team at Cisco Systems. It recommends reviewing no more than 200 to 400 lines of code at a time, and reports that the proportion of defects found falls away when a reviewer goes faster than about 500 lines an hour. That is one vendor’s study of one team, so treat the figures as a rule of thumb. The principle behind them is widely shared: Google’s published engineering practices also advise small changes, on the grounds that they are reviewed more quickly and more thoroughly.

An earlier version of this article repeated figures for the errors and money saved by inspections at well-known organisations. We could not trace them to a source we could read, so they have been removed.

What a reviewer looks for

Google’s guidance for reviewers lists design, functionality, complexity, tests, naming, comments, style and documentation. In terms of a business system, the questions are these:

  • Does the change do what was asked, including the awkward cases?
  • Is it more complicated than it needs to be? Will the next developer be able to follow it?
  • Does it come with tests?
  • Does it handle user input and permissions safely, and keep passwords and keys out of the code?
  • Does it fit the way the rest of the system is built?
  • If it alters the database, what happens to the existing data, and can the alteration be reversed?

The same guidance sets a sensible standard for approval: a change should be accepted once it clearly improves the overall health of the code, even if it is not perfect. Reviews that hold out for perfection stall the work and sour the team.

What makes reviews work, and what makes them fail

Small changes. A reviewer given a thousand lines skims them and approves. A reviewer given two hundred reads them.

Machines do the dull part. Formatting and style are checked by tools, so the reviewer’s attention goes on substance.

Reviews are prompt. A change that waits days for review holds up everything behind it, and the author has moved on to other work by the time the comments arrive.

Comments are about the code. A team where people defend “my code” learns less from review than one where the code belongs to everyone.

The common failure is the rubber stamp: approvals given in seconds because the reviewer is busy and trusts the author. The record shows that every change was reviewed, and none of them was read.

Review is not a substitute for testing. A reviewer reads the code and reasons about it. A test runs it. The kinds of test, and what each catches, are set out in our overview of software testing strategies.

Why this matters to the person paying

The system outlives the developer. If every change has been read by a second person, at least two people understand each part. When one leaves, the knowledge does not go with them. Where a system has had a single author for years, the developer leaving is a serious event for the business.

A sole developer has nobody to review their work. This is one of the practical differences between a software company and a freelance developer. If you rely on one person, consider paying for a periodic outside review.

The record is documentation. A pull request shows what was changed, the reason and who agreed. Years later, that history explains decisions nobody remembers making.

Unreviewed shortcuts accumulate. Review is where a quick fix is questioned before it becomes permanent. Without it, technical debt builds up unnoticed.

Review carries extra weight on an inherited system with no automated tests. There, a reviewer asking “what else uses this?” is one of the few safeguards available, alongside rehearsing the change on a copy of the system and having a named person in the business confirm the workflow still behaves correctly.

A code audit is a related but different thing. A code review examines one change. An audit examines the whole system at a point in time and reports on its condition, including how hard the code is to change safely.

Questions to ask your supplier or team

  • Is every change reviewed by a second developer before release, and does the tool enforce it?
  • Who reviews changes in an area only one developer knows?
  • What do the automated checks cover before a person looks?
  • How large is a typical change?
  • Do we own the repository, and can we see the review history?

Then ask to be shown one recent pull request. Its size, the comments on it and the time between submission and approval will tell you most of what you need to know about how seriously review is taken.

Tell us about your system

Say what it does, what it is built on and what is worrying you. We will reply with what we would look at first and whether we are the right people to help.

Tell us about your system 0800 433 7990 Monday to Friday, 9am to 5pm. A first 20-minute call is free, and we reply to every enquiry within one working day. What happens after you get in touch