Skip to main content

Data

Data Quality Is a Process Problem

You cannot clean your way out of a data quality problem that is created fresh every day by the process upstream of it.

8 min read

Every organisation that invests in a data warehouse eventually discovers the same uncomfortable fact: the warehouse does not fix data quality, it just gives bad data a faster and more visible way to reach everyone. Dashboards that used to be wrong quietly are now wrong publicly, in real time, in front of an executive team that assumed the investment in data infrastructure had already solved this.

The instinct at that point is to add more cleaning: validation rules, deduplication jobs, reconciliation scripts that run overnight and patch the worst of the damage before anyone in leadership sees it. This treats the symptom while leaving the disease untouched. Data quality problems are, in the overwhelming majority of cases we see, not created in the warehouse at all. They are created upstream, in the operational process that generates the data in the first place, and no amount of downstream cleaning permanently fixes a problem that is manufactured fresh every single day.

Where bad data is actually born

Consider a customer address that is wrong in a reporting system. The instinctive response is to fix the address in the data warehouse. But the address was wrong because a call centre agent, working against a script optimised for handle time rather than accuracy, typed it in incorrectly at the point of capture, and nothing about that process has changed. The warehouse fix corrects yesterday's error. Tomorrow's identical error is already being created, by the same process, for the same structural reason, and will need fixing again next week.

This pattern repeats across nearly every data quality issue worth solving: duplicate customer records created because two systems capture the same customer through different channels with no shared identifier, inconsistent product categorisation because two departments describe the same item differently for their own convenience, missing values because a form makes a field optional that the business actually depends on being filled in. In every case, the data is a faithful record of a flawed process, not a flaw in the data itself.

A warehouse does not create data quality. It just gives bad data a faster route to the people making decisions with it.

PrimeReach Consulting

Ownership is the missing ingredient

The reason this problem persists in most organisations is that nobody owns data quality at the point where it is created. Data teams are held accountable for the state of data in the warehouse, which they did not generate and often cannot fix at the source, while the operational teams that actually create the data have no accountability for its quality at all, because their performance is measured on speed, volume or cost, never on the downstream accuracy of what they produce.

Fixing this requires a genuinely uncomfortable conversation about incentives. A call centre measured purely on average handle time will always sacrifice data accuracy for speed when the two are in tension, because that is what the metric rewards. Until data quality becomes a measured, owned responsibility of the operational team that creates the data, rather than an unfunded mandate handed to a downstream data team, the same errors will keep arriving at the same rate, indefinitely.

  • Trace data quality issues back to the operational process that created them
  • Assign ownership of quality to the team generating the data, not only the team storing it
  • Change the incentive if the process rewards speed at the expense of accuracy
  • Treat downstream cleaning as containment, not as a cure

What good ownership looks like in practice

In organisations that manage this well, a data quality issue triggers a conversation with the process owner upstream, not just a fix in a transformation pipeline. Metrics exist at the point of capture, not only at the point of reporting, so a call centre or a sales team can see the accuracy of what they produce and is measured on it alongside their existing targets. Systems are redesigned, where possible, to make the correct behaviour the easy behaviour: shared identifiers that prevent duplicate records rather than deduplication scripts that clean them up afterwards, mandatory fields that reflect what the business actually needs rather than what was convenient to build.

None of this eliminates the need for data cleaning entirely; some quality issues are genuinely historical and some systems cannot be changed quickly. But it shifts the balance of effort from a permanent downstream cleanup operation towards a smaller, shrinking one, because fewer new errors are being created at the source every day.

Why this rarely gets fixed anyway

The honest reason most organisations never make this shift is that it requires data leaders to have influence over operational processes they do not own, and operational leaders to accept a new metric they did not ask for. Both of those are organisational fights, not technical ones, and they are considerably less comfortable to have than commissioning another cleaning tool.

But the maths is unforgiving. A cleaning pipeline that processes the same category of error every week, indefinitely, is not solving a problem, it is subsidising one. The organisations that eventually break the cycle are the ones willing to have the harder conversation about where the data actually comes from, and to hold the right team accountable for what it produces.

If your organisation is still running the same reconciliation job every week to fix the same category of error, the tool is not the problem. The process manufacturing that error every day has never been asked to change, and it will not change on its own.

Wherever you are in your transformation journey, let鈥檚 define the next move.

Start a conversation

Wherever you are in your transformation journey, let鈥檚 define the next move.

Start a conversation