Dabish Digital
Architecture

Event-driven architecture: a practical guide

It comes up on almost every project, usually later than it should. This guide covers what event-driven architecture actually involves, where it usually goes wrong, and how to tell whether yours is in reasonable shape.

The right architecture for a team of three is the wrong one for a team of thirty, and vice versa. Nothing below assumes a large team or a large budget — most of it is a decision somebody has to make and then write down.

What it costs to ignore

Events decouple producers from consumers, which is the whole point. None of that requires a large budget, only a decision and someone to own it. Check it against what you would want a competitor's site to get wrong.

For most businesses the question is not whether this matters but how much of it is worth doing right now. That depends on what you are trying to achieve in the next few months, not on best practice in the abstract. The teams that handle this well are rarely the ones with the biggest budgets.

The practical version

Debugging gets harder the moment nothing is synchronous. There is a version of this that is over-engineered, and it is worth avoiding. Anything you cannot measure here, you are deciding by taste, which is fine as long as everyone knows it.

Architecture is the set of decisions that are expensive to reverse, which is the only reason they deserve the name. The version that works in practice is usually less elaborate than the version described in the guides.

Start with one event, not with an event bus. Small and consistent beats large and occasional here. It rarely shows up as a line item, which is exactly why it slips.

A working checklist

If you want a quick read on where you stand, work through this. Anything you cannot answer confidently is where to start.

  • Events decouple producers from consumers, which is the whole point
  • Debugging gets harder the moment nothing is synchronous
  • Start with one event, not with an event bus
  • Someone is named as the owner, not just assumed to be
  • There is a date in the calendar to review it again
  • The decision and the reasoning behind it are written down somewhere findable
  • You could explain the current setup to a new hire in five minutes

Where it usually goes wrong

The most common failure is not doing this badly. It is doing it once, during a launch, and never revisiting it. Circumstances move, the setup does not, and the gap widens quietly until something breaks or somebody notices the numbers.

  • It was configured during a launch and has not been touched since
  • Different people in the business believe different things are true about it
  • There is no way to tell whether the last change helped or hurt
  • The only person who understands it has left, or is about to

Most systems fail at the seams rather than inside any one component. The teams that handle this well are rarely the ones with the biggest budgets.

How we approach it

On our projects this gets handled during the build rather than added afterwards, because retrofitting it costs several times more than including it. We write down what was decided and why, so the next person to touch it is not guessing.

If you are working with someone else, the questions worth asking are simple: who owns this, how will we know it is working, and what happens when it needs to change?

A reasonable first step

Pick the single item from the checklist above that would cause the most trouble if it turned out to be wrong. Fix that one, confirm it worked, then move on. If you want a second opinion on how yours is set up, ask.