DevOps & Engineering

Site Reliability Engineering (SRE)

Service level objectives, error budgets and an on-call practice that does not burn out your team.

Overview

Reliability work goes wrong when it becomes an aspiration to be as available as possible. That target is unbounded and unaffordable. SRE replaces it with a number the business agrees to: this service will meet this objective, and here is what we do when it does not.

We help teams define service level objectives that reflect what users actually notice, instrument them, and use the resulting error budget to make the reliability-versus-features trade-off explicit rather than argued.

What you get

Objectives users recognise

SLOs based on what customers experience, not on infrastructure metrics that look reassuring during an outage.

Alerts worth waking for

Paging only on symptoms that need a human now; everything else becomes a ticket or a dashboard.

Sustainable on-call

A rotation, handover and escalation model designed for a small team rather than borrowed from a large one.

Incidents that teach

Blameless reviews producing a small number of changes that actually get made.

How we work

  1. 01

    Define

    Critical user journeys identified and translated into measurable objectives with the business.

  2. 02

    Instrument

    Metrics and traces implemented so the objectives can be measured honestly and continuously.

  3. 03

    Respond

    Alert routing, on-call rotation, escalation and runbooks put in place and rehearsed.

  4. 04

    Review

    A regular reliability review where the error budget drives what gets prioritised next.

Common questions

We are ten people. Is SRE overkill?

The full Google model is. Defining two or three SLOs and cutting noisy alerts is valuable at almost any size, and takes days rather than months.

What availability target should we set?

Lower than instinct suggests. Each additional nine multiplies cost and engineering effort. The right number is the one the business will genuinely pay for.

Can you run on-call for us?

We can provide support cover as part of a managed arrangement, though for product-specific incidents your engineers will always be the faster responders.

Often paired with

Ready to talk about site reliability engineering (sre)?

We will tell you what we would do, roughly what it costs, and whether it is worth doing yet.

Book a meeting