Dabish Digital
Cloud

Getting started with alerting

Every audit we run turns up some version of this. A short on-ramp to alerting for teams who have not touched it before.

Operability is a feature, and it has to be built rather than bought. It is the sort of thing that looks like polish right up until it costs you an enquiry.

The reason this keeps coming up

An alert nobody acts on trains everyone to ignore alerts. It is worth being explicit about, because assumptions differ quietly. If two people in the business would answer this differently, that gap is the actual problem.

Your first week

  1. Find out what is already in place
  2. Page on symptoms customers feel, not on every anomaly
  3. Change one thing and measure it

Every alert should link to what to do about it. Where this goes wrong is almost never a lack of knowledge. The practical test is whether someone new to the project could tell, in a minute, that it had been handled.

What this looks like day to day

Cloud work rewards teams who automate early and punishes teams who click through consoles. Three things worth confirming about alerting before you move on:

  • Someone can say what the current setup is without going to look
  • Every alert should link to what to do about it — and you know whether that is true here
  • There is a way to tell whether the last change to this helped

Most of the value here comes from doing the first two things, not all of them.