Signs it is time to revisit alerting
The version of this that works is simpler than the version most people imagine. A few signals that alerting is due some attention.
The bill is a design document: it tells you exactly what your architecture actually does. If two people in the business would answer this differently, that gap is the actual problem.
The signals
- Nobody can say when it was last reviewed
- The answer depends on who you ask
- An alert nobody acts on trains everyone to ignore alerts
- Page on symptoms customers feel, not on every anomaly
Where to start
Every alert should link to what to do about it. None of that requires a large budget, only a decision and someone to own it. The failure mode is not doing it wrong, it is doing it once and assuming it stays done.
The short version
Cloud work rewards teams who automate early and punishes teams who click through consoles. Three things worth confirming about alerting before you move on:
- Someone can say what the current setup is without going to look
- Every alert should link to what to do about it — and you know whether that is true here
- There is a way to tell whether the last change to this helped
If you want a second opinion on how yours is set up, ask.