Dabish Digital
Cloud

A short guide to observability

We end up explaining this on discovery calls often enough that it deserved writing down. Everything we would tell a client about observability in the time it takes to drink a coffee.

Cloud work rewards teams who automate early and punishes teams who click through consoles. The version that survives contact with a real deadline is the simple one.

The reason this keeps coming up

Monitoring tells you something broke, observability tells you why. There is a version of this that is over-engineered, and it is worth avoiding. It is worth deciding this deliberately rather than inheriting whatever the last person set up.

What good looks like

Traces, metrics, and logs answer different questions. Getting it slightly wrong is survivable. Ignoring it entirely is not. Check it against what you would want a competitor's site to get wrong.

Warning signs

Instrument the paths that lose money first. It is worth being explicit about, because assumptions differ quietly. The teams that stay on top of it are the ones who put it on a calendar rather than a wish list.

In practice

Operability is a feature, and it has to be built rather than bought. Three things worth confirming about observability before you move on:

  • Someone can say what the current setup is without going to look
  • Traces, metrics, and logs answer different questions — and you know whether that is true here
  • There is a way to tell whether the last change to this helped

If you are not sure where your systems currently stand on this, it takes us about an hour to find out.