A software testing strategy is a decision about which kinds of test to run, at which level of the system, how often and by whom, so that the effort goes where a failure would hurt most. Almost every strategy is assembled from the same four levels of testing: unit, integration, system and acceptance. On top of those sit a few checks on qualities such as speed and security. This overview explains what each one catches, what it misses and how they are combined.
The four levels of testing
| Level | What it checks | Usually done by | How often |
|---|---|---|---|
| Unit | One rule or calculation on its own | Developers, automated | Every change |
| Integration | Parts working together, such as the application and its database | Developers, automated | Every change, or daily |
| System | Whole workflows through the real screens | Testers, partly automated | Before each release |
| Acceptance | Whether the software does the job the business needs | The people who will use it | During the build and before go-live |
Unit testing takes the smallest piece of code that does one thing and checks it in isolation. If a password must be at least eight characters, the unit tests try seven characters, eight, nine and an empty value. Unit tests run in seconds and, when one fails, it points at the exact rule that broke. What they cannot tell you is whether the pieces fit together.
Integration testing checks the joins. One test saves an order to the database and confirms it comes back with the same totals. Another confirms that the export file loads into the accounts package. In our experience many of the faults in business systems live at these joins: a date in the wrong format, or a field one side treats as optional and the other requires. Integration tests are slower than unit tests and need a test database or a stand-in for the outside system.
System testing, often called end-to-end testing, exercises the whole application as it will be deployed. A test places an order, raises the invoice and records the payment, through the same screens a user would. It can be done by hand or automated with a browser tool such as Selenium. It finds faults in configuration and wiring that no smaller test can see. These tests are slow and break when screens are redesigned, so a sensible strategy keeps them few.
Acceptance testing asks a different question. The first three levels ask whether the software works as specified. Acceptance asks whether what was specified is what the business needs. Take a requirement such as “the system alerts a manager when someone tries to log in after 7pm”. It looks clear until somebody asks which manager, alerted how, and 7pm in which time zone. Putting the software in front of the people who will use it is how such gaps come to light, and it works best in small pieces throughout a project.
Tests for qualities as well as features
Some of the most damaging faults are not wrong answers. The answer is right and arrives too slowly, or reaches the wrong person.
- Performance. A screen that is quick with a hundred test records can be unusable with five years of real ones. Test with realistic volumes of data and a realistic number of users.
- Security. Automated tests can confirm that one user cannot see another’s records, and scanning tools can flag components with known vulnerabilities. A penetration test, where specialists try to break in, is a separate service bought from a specialist firm.
- Recovery. Restore a backup into a spare environment and time it. A backup that has never been restored is an assumption.
- Compatibility. For anything used through a browser, check the browsers and devices your users have.
Regression testing is not a fifth level. It is the practice of re-running tests from every level after a change, to confirm that what worked before still works.
How the levels fit together
The usual guide is the test pyramid, which Mike Cohn described in his book Succeeding with Agile. Picture the levels stacked with unit tests at the bottom. The pyramid says to have many small, fast tests at the base and progressively fewer as you go up.
The reasoning is about cost. Unit tests are cheap to run and precise about what failed. End-to-end tests are slow and fragile. A suite made only of end-to-end tests takes hours to run and fails for reasons unrelated to real faults, until the team stops trusting it. A suite made only of unit tests can pass in full while the application fails to start.
At any level, a test can be designed in one of two ways: from the outside, knowing only what the system should do, or from the inside, knowing how the code works. That distinction is the subject of our article on black box and white box testing.
Choosing a strategy for your system
There is no single correct mix. Start by listing where a failure would cost the most, and weight the testing accordingly.
- A system that calculates prices, pay or tax needs thorough unit tests on the calculation rules, and acceptance by whoever in finance knows the right answers.
- A system whose main job is moving data between other systems needs most of its effort at the integration level.
- A customer-facing portal needs end-to-end tests of the main customer tasks, plus attention to performance and security.
- A small internal tool with a short life ahead of it may need little beyond a written checklist.
Whatever the mix, it should fit on a page: what is tested at each level, what is automated, who accepts the work, and what is not tested. The last item is the most informative. Every strategy leaves something out, and it is better to know what.
When the system has no tests at all
The pyramid assumes tests are written as the code is written. Many older business systems have none, and often no description of what they should do. Whether that is true of yours is one of the things a code audit reports.
Such a system cannot be tested from the bottom up. Old code is frequently arranged so that single rules cannot be run in isolation, and nobody is sure what the rules are. So the order is reversed. This is how we approach it when taking over a system someone else built:
- Record what the whole system does today for real inputs, such as last month’s orders and the invoices they produced. These recordings are called characterisation tests.
- For each workflow that matters, name a person in the business who can confirm the recorded behaviour is correct.
- Rehearse every change on a restored copy and compare the results with the recording.
- Each time an area of code is changed, add unit tests around it.
Over time the pyramid fills in from the top down, concentrated on the parts of the system that change.
Questions that reveal the strategy
- What is tested automatically on every change?
- What does a person check before each release, and from what list?
- Who in our business accepts the work, and at what point?
- Has anyone tested with a realistic volume of data?
- When was a backup last restored as a test?
- What is not tested?
Ask whoever looks after your system for the one-page answer. If it exists, you will learn where the gaps are. If nobody can produce it, there is no strategy, and that is worth knowing before the next change goes live.