Data quality before AI — why pilots fail on master data
An AI pilot promises quick wins and is still stuck months later. The cause is rarely the model, but the data beneath it: inconsistent keys, duplicates, missing units. This article describes why data maturity is the actual project — and when AI honestly does not help.
A project starts with a clear promise: AI is to speed up a laborious process in pharma, MedTech or care delivery. After the first few weeks it stalls. What then gets discussed is usually the model — while the problem almost always sits one level below, in the data the model is meant to work with.
Why the data comes first
A model is only as good as the data it learns from and receives in operation. That sounds obvious, yet it is regularly skipped during planning, because attention goes to the visible part: the prediction, the automation, the interface. The invisible part — master data, consistency, completeness — is what decides whether anything dependable comes out at all.
Master data means the stable core records a process draws on again and again: articles, customers, locations, units. Where these are maintained inconsistently, even an excellent method cannot close the gaps — it merely passes them on.
A pattern from practice
A recurring picture, in our experience: the same item carries three different numbers across three systems. A customer exists twice, once with and once without an umlaut. Quantities appear sometimes as units, sometimes as packs, without the unit being carried along. To a person these are trivia they reconcile in their head. An automated method cannot — it treats two spellings as two things and calculates on a basis that does not hold. The pilot then delivers results nobody wants to stand behind, and loses its credibility before it could show its worth.
How to proceed: data maturity before the model
We reverse the usual order and begin with the data:
- Inventory the data sources: which systems supply what, in what form, how current?
- Clean up keys and master data: unique identifiers, consistent units, duplicates merged.
- Set one small, demonstrable metric instead of a grand vision — something that can be measured honestly after four weeks.
- Only then the model. It is the last step, not the first.
This route feels slower, because the effort becomes visible at the beginning instead of at the end. In total it is shorter, because the rework falls away.
Where this does not apply
Not every problem is a data problem, and not every data problem needs AI. Where the volume of data is too small, where the process changes faster than data can be gathered, or where the bottleneck is organisational rather than analytical, even the best model will not help. In many cases a clear rule or a good dashboard beats a method nobody can follow. And a broken process should not be automated but put in order first — otherwise you scale the error.
Conclusion
Data maturity is not the preliminary stage of the project, it is the project. The model is the last twenty per cent, the part that is visible — what carries it is the work beforehand. Start there honestly, and you reach results you can rely on with less effort.
We begin with a process analysis and show, with evidence, what is possible.
Arrange a conversation