Data quality: the questions we get asked most
It comes up on almost every project, usually later than it should. The questions about data quality that come up most often on our calls.
Most data problems are ownership problems that turned into technical ones. Budget a little time for it every quarter and it never becomes a project of its own.
Do we need to care about this?
Bad data spreads faster than anyone corrects it. There is a version of this that is over-engineered, and it is worth avoiding. The practical test is whether someone new to the project could tell, in a minute, that it had been handled.
Can it wait until after launch?
Occasionally. More often the post-launch version costs several times the pre-launch one. Getting it slightly wrong is survivable. Ignoring it entirely is not.
How do we know it is working?
Someone must own each dataset or nobody does. That sounds obvious written down. It is still the thing most often skipped. If two people in the business would answer this differently, that gap is the actual problem.
The short version
Numbers get quoted in meetings long after anyone remembers how they were calculated. Three things worth confirming about data quality before you move on:
- Someone can say what the current setup is without going to look
- Someone must own each dataset or nobody does — and you know whether that is true here
- There is a way to tell whether the last change to this helped
Pick the one that would hurt most if it failed, and start there.