Dabish Digital
Cloud

A short guide to alerting

Teams tend to reach for this after something has already gone wrong. Everything we would tell a client about alerting in the time it takes to drink a coffee.

Operability is a feature, and it has to be built rather than bought. The version that survives contact with a real deadline is the simple one.

The reason this keeps coming up

An alert nobody acts on trains everyone to ignore alerts. Small and consistent beats large and occasional here. Budget a little time for it every quarter and it never becomes a project of its own.

The practical version

Page on symptoms customers feel, not on every anomaly. This is the sort of thing that compounds, quietly, in both directions. The failure mode is not doing it wrong, it is doing it once and assuming it stays done.

Warning signs

Every alert should link to what to do about it. The reasoning matters more than the rule, because the rule has exceptions. If two people in the business would answer this differently, that gap is the actual problem.

The short version

Cloud work rewards teams who automate early and punishes teams who click through consoles. Three things worth confirming about alerting before you move on:

  • Someone can say what the current setup is without going to look
  • Page on symptoms customers feel, not on every anomaly — and you know whether that is true here
  • There is a way to tell whether the last change to this helped

If any of that sounds like a description of your current setup, it is fixable.