The difference between black box and white box testing is what the tester knows when designing the test. In black box testing the software is treated as a sealed unit: the tester supplies inputs, checks the results against what should happen and never looks at the code. In white box testing the code is open, and tests are written to exercise particular lines, branches and paths inside it. Each finds faults the other cannot, which is why a well-tested system has both.
You will also see the names closed box and opaque box for the first, and clear box, glass box or open box for the second. They mean the same things.
Black box testing: judging by behaviour
A black box test starts from a statement of what the system should do. Suppose the rule is that orders of £500 or more get five per cent off. A black box tester does not care how that is programmed. They choose inputs and work out the expected answer for each:
- an order of £499.99, which should get no discount;
- an order of exactly £500.00, which should;
- an order of £500.01;
- an order of zero, a negative amount and an order with a missing price, all of which should be rejected cleanly.
Two techniques are at work here. Equivalence partitioning means grouping the inputs that ought to behave alike and testing one from each group, so you do not test a thousand orders that are all plainly over the threshold. Boundary value analysis means testing at the edges of each group, because that is where mistakes collect: a programmer who wrote “more than” where the rule said “or more” is caught by the £500.00 case and by no other.
Black box tests can be run by hand or automated, and they do not need a programmer to design them. They need someone who knows what the right answer is. That makes them the natural form for acceptance testing, where the people who will use the system confirm it does the job.
Their limit is that they only reach the code the chosen inputs happen to pass through. If the program contains a special case for one customer, and no test order belongs to that customer, the special case goes unchecked.
White box testing: judging by the code
A white box tester reads the code first. Reading the discount routine, they might find that staff accounts get a different rate, and that orders dated before a certain day follow an older rule. Nobody working from the outside would know to try those. The white box tester writes a test for each branch.
The usual measure is code coverage. A tool records which lines and branches of the code ran while the tests were running and reports the proportion. Its real use is to show what was never run. It is a weaker guide to quality than it looks, because a test can pass through a line without checking that the line produced the right result.
Unit tests, the small tests developers write alongside the code, are usually designed this way, since the developer has the code in front of them. A code review, where a second developer reads a change before it is accepted, applies the same inside view without running anything.
White box testing is good at finding untested branches, faulty error handling and weaknesses that are only visible in the code, such as a database query assembled from text a user typed in. Its limit is the mirror image of the other kind: it can only examine what was written. If a requirement was forgotten, there is no code to read, and white box tests will confirm that the program does what the program says.
The two side by side
| Black box | White box | |
|---|---|---|
| What the tester knows | What the system should do | How the code does it |
| Tests are derived from | Requirements, workflows, past behaviour | Branches, paths and conditions in the code |
| Usually done by | Testers and business users | Developers |
| Most common at | System and acceptance testing | Unit testing |
| Good at finding | Missing or wrong behaviour | Hidden special cases, untested paths |
| Blind to | Code the chosen inputs never reach | Requirements nobody implemented |
| When the code is restructured | Still valid if behaviour is unchanged | Often needs rewriting |
In practice a lot of testing sits between the two and is called grey box testing. The tester works from the outside but uses some knowledge of the inside, for example checking the database after submitting a form to confirm the right rows were written. Our overview of software testing strategies covers the levels of testing where each approach is used.
The same words appear in security work. A penetration test, where specialists attempt to break into a system, may be offered as black box or white box. The National Cyber Security Centre’s guidance on penetration testing draws the same line: in the first the testers are told nothing about the internals of the target, and in the second they are given full information. It is a separate discipline from the functional testing described here, carried out by specialist firms.
Which to use on a system nobody fully understands
The textbook order is white box tests while the code is written and black box tests afterwards. An inherited system turns that round. When we take over software that someone else built, there are often no tests and no specification, and the code cannot be trusted as a statement of what the business wants. It only says what the system does.
So the first tests are black box. A characterisation test records the system’s present output for a set of real inputs, such as last month’s orders and the invoices they produced. Someone in the business who uses that workflow is named as the person who confirms the recorded behaviour is correct. The recording then becomes the baseline for regression testing: after any change, the same inputs must give the same outputs unless a difference was intended.
White box work comes second, and it has a specific job. A coverage tool shows which parts of the code the recorded cases never reached. Those parts are usually the special cases: the customer on different terms, the routine that runs once a year. Each becomes a question for the named person. “The code treats accounts flagged this way differently. Is that still wanted?” The answer produces either a new test or a piece of code that can be retired.
All of this is rehearsed on a restored copy of the system, never on the live one.
The last row of the table matters most when a system is being modernised. Black box tests do not depend on how the code is arranged, so tests recorded against the old system can be run against its replacement. If the same inputs give the same outputs, the rules have survived the move.
What to check with your supplier
- Ask how your system is tested. A good answer mentions both kinds: tests developers write against the code, and checks made from the user’s side.
- If you are given a coverage percentage, ask which parts of the system fall outside it. What matters is whether the calculations, permissions and integrations are covered.
- Ask who in your business carries out acceptance testing. This is black box testing only you can do, because only you know what a correct result looks like.
- If you are buying a penetration test, ask whether it is black box or white box, and why that suits your situation.
A quick way to find out where you stand: pick one rule your business depends on, such as a price calculation, and ask to see the tests for it. The answer will tell you whether anyone has checked it from the outside, from the inside, or at all.