Three myths about data quality
It is rarely the thing that gets a project approved, and often the thing that decides how it goes. A few things about data quality that get repeated more often than they get checked.
Numbers get quoted in meetings long after anyone remembers how they were calculated. If two people in the business would answer this differently, that gap is the actual problem.
“It only matters for big sites”
Bad data spreads faster than anyone corrects it. None of that requires a large budget, only a decision and someone to own it. It is worth deciding this deliberately rather than inheriting whatever the last person set up.
“We can deal with it after launch”
Sometimes true, usually expensive. Getting it slightly wrong is survivable. Ignoring it entirely is not.
“Our platform handles it”
Someone must own each dataset or nobody does. There is a version of this that is over-engineered, and it is worth avoiding. The failure mode is not doing it wrong, it is doing it once and assuming it stays done.
In practice
Most data problems are ownership problems that turned into technical ones. Three things worth confirming about data quality before you move on:
- Someone can say what the current setup is without going to look
- Someone must own each dataset or nobody does — and you know whether that is true here
- There is a way to tell whether the last change to this helped
Pick the one that would hurt most if it failed, and start there.