A service level agreement, or SLA, is the part of a support contract that says how quickly the supplier will respond to a problem, during which hours, and what happens if it does not. It is a good idea to have one for any software the business depends on, because without it the word “urgent” has no agreed meaning and you have nothing to measure the service against.
This article explains what an SLA does for you, what it cannot do, and how to tell a useful one from a page of comfortable wording.
What a service level agreement covers
An SLA belongs to an ongoing service. For software, that means support and maintenance after a system has gone live. A project to build something new is governed by different things: a scope, a price and acceptance checks.
A software support SLA normally sets out the hours of cover, a small number of severity levels, a response time for each level, the route for escalating a problem and how performance is reported. We list everything a contract should contain, including the SLA, in what a software support contract should include.
It helps to separate two kinds of SLA that are often confused. A hosting provider’s SLA promises availability: the servers will be running for a stated proportion of the time. A support SLA promises attention: a qualified person will start work on your problem within a stated time. A bespoke system usually needs both, from different parties, and one does not substitute for the other.
Why it is worth having
Priorities are settled in advance. The worst moment to negotiate how serious a fault is, and how fast someone should act, is during the fault. An SLA makes that decision calmly, months earlier, with both sides in the room.
You can measure the service. With targets written down and tickets timed against them, a review of the supplier rests on a report. Without them it rests on how the last conversation went.
It tells you what you are paying for. A supplier that commits to respond within hours has to keep people available, and that costs more than a commitment to respond next day. Seeing the levels side by side lets you buy the cover the business needs and no more.
It protects the supplier as well. The agreement states what is covered, in which hours, and what counts as separate work. A supplier with clear limits can plan its staffing, and a supplier that can plan gives a steadier service.
It outlasts the people. Informal arrangements work while the two people who made them are still in post. When either moves on, the understanding goes with them. A written agreement stays.
What an SLA cannot do
An SLA is a promise about behaviour. It is not a guarantee about the software.
A response time is not a fix time. Most suppliers of bespoke software commit to how soon work starts, because how long a repair takes depends on a cause nobody yet knows. Be wary of a promise to fix any fault within a fixed number of hours. Either it is hedged somewhere else in the document, or it will be met with hurried patches.
Service credits rarely cover the loss. Many SLAs give a credit against the fee when a target is missed. The credit is a signal and an incentive. It will be small beside the cost of a day without your order system. Some contracts also state that credits are your only remedy for poor service. Whether an SLA is enforceable, and what you can claim if it is breached, depends on how the contract is written, so have a solicitor read it before you sign.
It does not make a neglected system reliable. If the software is on an out-of-support platform, has no tested backup and can only be released by one person, quick responses will not prevent the outages. The planned work matters as much as the reactive work, which is why we treat support and maintenance as two halves of one job.
It means little from a supplier who does not know the system. A team that has never built your application from its source code cannot honestly promise how it will handle a fault. This is why we set the terms of an agreement after assessing the system, not before.
Five things to check in a draft
-
How severity is defined. Each level should be described by its effect on the business: nobody can work, one function has failed but there is a way round it, or the fault is cosmetic. Check who decides the level when you and the supplier disagree.
-
What “response” means. An automatic email saying the ticket has been received is not a response. Look for wording that requires a person able to work on the problem to have started.
-
When the clock runs. A four-hour response within working hours means a fault reported at five in the afternoon may be picked up the next morning. Check the hours of cover, the treatment of bank holidays, and whether the clock stops while the supplier waits for information from you.
-
What happens when a target is missed. Look for a named person to escalate to, and for a consequence of repeated misses, such as the right to end the contract early.
-
How performance is reported. You should receive a regular report showing each ticket, its severity and whether the target was met, drawn from the ticket system and not compiled from memory. Our guide to software support ticket systems covers what to expect from one.
Choosing the right level of cover
More cover is not always better. It costs more, and the money may do more good spent on preventing faults.
Work out what an hour without the system costs, and what a day costs. A system used by six people in one office to produce monthly reports can wait until the next working day. A system that takes customer orders, runs a production line or pays people is a different matter, and may justify cover outside office hours.
Then look at when the system is used. Cover from nine to five is no help to a warehouse that starts at six.
Suppliers usually publish a few levels of cover so that you can match one to the system. Ours are set out under the support plans on our software maintenance page.
Making it work after you sign
An SLA that nobody looks at decays. A few habits keep it useful.
Report every problem through the agreed route. A telephone call to a helpful developer gets the fault fixed, but it leaves no record, and the report at the end of the month will show a service that was never tested.
Read the report when it arrives, and hold a short review with the supplier every quarter. Ask which faults recurred and what is being done about the cause.
Try the urgent route once before you need it. Find out whether the out-of-hours number is answered.
Before you sign a draft, take the last three serious faults the system suffered and ask the supplier how each would have been classed and handled under the agreement. The answers will show whether the document describes a service you would have been happy with.