Webneuron
Managed Services

Your SLA is not a reliability strategy

An SLA describes what happens after something breaks and who pays. It says almost nothing about whether it will break.

June 23, 20267 min readBy Webneuron Engineering Team

Service level agreements are a commercial instrument. They define availability targets, response and resolution windows, escalation paths, and what compensation is due when a target is missed. They are useful, negotiable, and entirely necessary in any serious managed services relationship.

What they are not is a plan for reliability. An SLA allocates the consequences of failure. It does not reduce the probability of failure, and organisations regularly conflate the two — signing a demanding agreement and treating the signature as the work.

What the numbers conceal

Availability targets are the clearest example. Three nines sounds rigorous until you notice that it permits roughly eight hours of downtime a year, and that whether those eight hours fall on a quiet Sunday or across a month-end close makes an enormous difference to the business and none at all to the measurement.

Response time is measured similarly loosely. A fifteen-minute response commitment is often satisfied by an acknowledgement. What matters is time to restoration, and that is governed by whether the responding engineer understands the system — which no clause can specify.

Most SLAs also measure the component rather than the journey. Every service can report itself healthy while a customer cannot complete a purchase, because the failure lives in the interaction between components and nobody owns the interaction.

Questions worth more than the clause

  • Who is actually on the rota, and have they operated this system before, or will they be reading a runbook for the first time during your incident?
  • What proportion of incidents are detected by monitoring rather than reported by users? This single figure describes the maturity of the arrangement better than any target.
  • What happens between incidents? A provider whose economics depend on minimal engagement will restore service and not address causes.
  • How are recurring incidents handled? The same failure three times is a design problem, and the contract should create an incentive to fix it rather than to keep responding.
  • What does the post-incident process produce, and does anything ever change as a result?

The incentive worth designing

The structural weakness of a conventional SLA is that it pays for response. A provider is compensated for being available when things break, which is not the same as being compensated for things not breaking.

The better arrangements introduce something that pulls in the other direction: shared reliability targets, an explicit allocation of effort to remediation rather than only to response, and a review cadence where recurring causes are named and scheduled. The clause still matters — but the reliability comes from the operating relationship, and no amount of contractual precision substitutes for a provider who understands your architecture before the pager goes off.

Let's build the system your business will run on next.

Tell us where it hurts. We'll bring the architects, engineers, and delivery model to fix it — and scale it.