Alerting: what to get right first
The version of this that works is simpler than the version most people imagine. If you only fix one thing about alerting this quarter, make it the first item below.
The bill is a design document: it tells you exactly what your architecture actually does. The version that survives contact with a real deadline is the simple one.
Start here
An alert nobody acts on trains everyone to ignore alerts. Where this goes wrong is almost never a lack of knowledge. It is the sort of thing that looks like polish right up until it costs you an enquiry.
Then this
Page on symptoms customers feel, not on every anomaly. There is a version of this that is over-engineered, and it is worth avoiding. The teams that stay on top of it are the ones who put it on a calendar rather than a wish list.
Eventually
Every alert should link to what to do about it. That sounds obvious written down. It is still the thing most often skipped. If two people in the business would answer this differently, that gap is the actual problem.
In practice
Operability is a feature, and it has to be built rather than bought. Three things worth confirming about alerting before you move on:
- Someone can say what the current setup is without going to look
- Every alert should link to what to do about it — and you know whether that is true here
- There is a way to tell whether the last change to this helped
The point is not perfection, it is knowing which of these you have consciously chosen to skip.