Dabish Digital
Cloud

Blue-green deployment: a practical guide

We end up explaining this on discovery calls often enough that it deserved writing down. This guide covers what blue-green deployment actually involves, where it usually goes wrong, and how to tell whether yours is in reasonable shape.

The bill is a design document: it tells you exactly what your architecture actually does. Nothing below assumes a large team or a large budget — most of it is a decision somebody has to make and then write down.

What it costs to ignore

Running two environments makes rollback instant. The reasoning matters more than the rule, because the rule has exceptions. If it only works because one person remembers to do something, it does not work yet.

For most businesses the question is not whether this matters but how much of it is worth doing right now. That depends on what you are trying to achieve in the next few months, not on best practice in the abstract. That sounds obvious written down. It is still the thing most often skipped.

What good looks like

Database changes are what make it complicated. It is worth being explicit about, because assumptions differ quietly. The practical test is whether someone new to the project could tell, in a minute, that it had been handled.

Operability is a feature, and it has to be built rather than bought. The version that works in practice is usually less elaborate than the version described in the guides.

Practise the rollback before you need it. There is a version of this that is over-engineered, and it is worth avoiding. The version that survives contact with a real deadline is the simple one.

A working checklist

If you want a quick read on where you stand, work through this. Anything you cannot answer confidently is where to start.

  • Running two environments makes rollback instant
  • Database changes are what make it complicated
  • Practise the rollback before you need it
  • Someone is named as the owner, not just assumed to be
  • There is a date in the calendar to review it again
  • The decision and the reasoning behind it are written down somewhere findable
  • You could explain the current setup to a new hire in five minutes

The mistakes we see most

The most common failure is not doing this badly. It is doing it once, during a launch, and never revisiting it. Circumstances move, the setup does not, and the gap widens quietly until something breaks or somebody notices the numbers.

  • It was configured during a launch and has not been touched since
  • Different people in the business believe different things are true about it
  • There is no way to tell whether the last change helped or hurt
  • The only person who understands it has left, or is about to

Cloud work rewards teams who automate early and punishes teams who click through consoles. This is the sort of thing that compounds, quietly, in both directions.

How we approach it

On our projects this gets handled during the build rather than added afterwards, because retrofitting it costs several times more than including it. We write down what was decided and why, so the next person to touch it is not guessing.

If you are working with someone else, the questions worth asking are simple: who owns this, how will we know it is working, and what happens when it needs to change?

Making it stick

Pick the single item from the checklist above that would cause the most trouble if it turned out to be wrong. Fix that one, confirm it worked, then move on. None of this needs a rewrite. Most of it is a morning's work once someone decides to do it.