A practical checklist for alerting
The version of this that works is simpler than the version most people imagine. Run through this the next time alerting comes up.
The bill is a design document: it tells you exactly what your architecture actually does. Budget a little time for it every quarter and it never becomes a project of its own.
The checklist
- An alert nobody acts on trains everyone to ignore alerts
- Page on symptoms customers feel, not on every anomaly
- Every alert should link to what to do about it
- Someone is named as the owner
- There is a date to review it again
What is actually at stake
An alert nobody acts on trains everyone to ignore alerts. The reasoning matters more than the rule, because the rule has exceptions. If two people in the business would answer this differently, that gap is the actual problem.
The short version
Operability is a feature, and it has to be built rather than bought. Three things worth confirming about alerting before you move on:
- Someone can say what the current setup is without going to look
- An alert nobody acts on trains everyone to ignore alerts — and you know whether that is true here
- There is a way to tell whether the last change to this helped
If you want a second opinion on how yours is set up, ask.