Why pharma data is fragmented — and what it really costs
Florian SeidlFlorian Werner
Florian Seidl, cts Group  ·  Florian Werner, Xenium AG
BU Manager Industrial Informatics & IT Principal Consultant · · 8 min read

Why pharma data is fragmented — and what it really costs

“Everybody needs data. Nobody speaks the same language.” That line captures the daily reality in many pharma plants better than any statistic. The problem is rarely missing data; it is missing shared context. Why that is the case — and where a data project really begins.

Walk into a pharma plant and look at the data, and you rarely see one system. You see many. Separate lines, individual package units, equipment from different vendors with different controllers. Each PLC (Programmable Logic Controller) speaks its own language. Unlike in chemicals, there is usually no central process control system (DCS, Distributed Control System) that ties everything together.

What remains is a grown, inconsistent data landscape — the logical result of how these plants came together over the years.

If you recognise this from your own plant: it looks that way in many of them. And it is the starting point that decides whether data ever becomes useful.

How the fragmentation happens

Equipment in pharma manufacturing is procured piece by piece, not as an integrated whole. Purchasing focuses on the process: the machine has to produce its product reliably. Data connectivity was rarely a criterion in the original design.

On top of that comes a long history. Many sites carry decades of technology, grown through individual investments. Over the years, local teams built pragmatic solutions that work for their line — but these island solutions do not scale across equipment, functions and sites.

The decisive point sits earlier than most people think: equipment is bought as a production asset, not as a data source. No one asks at purchase: what do we want to do with this data later? Data context gets bolted on afterwards instead of being designed in from the start.

The real problem is not that data is missing

The valuable data is already there. It sits in the MES (Manufacturing Execution System), the historian, the ERP (Enterprise Resource Planning), the LIMS (Laboratory Information Management System) and the automation systems. The problem is not its volume.

The problem is that it sits in silos, follows different structures and carries no shared meaning. Different sites use different naming conventions, standards and operating practices. The same business question then needs data from several disconnected systems.

And this is where expectations diverge. Manufacturing, Quality, IT and Management all agree that data matters — but not why it matters. Each department defines success differently. Each optimises for a different goal.

Everyone needs the same data — and yet everyone comes up empty. Not because the data is missing, but because no one speaks the same language. This is not a data problem. It is a problem of context and alignment.

Try the test for yourself: ask Manufacturing, Quality, IT and Management separately what “good use of data” concretely means to them. If you get four different answers, that is where the real project lies — not in the next interface.

Why more data does not help here

The obvious reaction is: just connect everything, then we will have the data. That falls short. Connectivity alone creates no value. Raw data without context is noise — only at scale.

Process values are available, but not linked to equipment, batch, process step or business meaning. The user gets more data, but not more insight. Worse still: data without context creates false confidence. You think you know what is happening — and you are wrong.

⚠ A pattern that repeats

Ten years ago the industry dumped everything into data lakes and looked for the value later. Back then it cost storage. Today the same pattern runs again — only now it burns compute and AI tokens instead of storage. Collect first, ask what for later: it was the wrong way round even then.

Where a data project really begins

Not at the level of individual data points, but with the business context, the data landscape and the use cases. Before anyone talks about a platform, there are four questions:

1
Which decisions or processes should improve? Not “what data do we have”, but “what do we want to change”.
2
Which business outcomes matter most? Opaque OEE (Overall Equipment Effectiveness), manual shift logs, slow quality reporting — the concrete pain points.
3
Which use cases create measurable value? One clear first use case beats a platform project without a goal.
4
Which data is actually relevant for that? Only once the why is settled is it worth looking at the what.

Only then comes the assessment of the data situation — along four dimensions. They decide whether a project can start yet, or whether a foundation is still missing.

Dimension 1

Availability

Which data exists — and which does not? Where does it come from, how fragmented is the landscape, how easy is it to access? Missing data creates blind spots.

Dimension 2

Quality

Completeness, consistency, reliability, traceability. Do “identical” data from different sources tell the same story? Can every value be traced back to its origin?

Dimension 3

Context

Which batch, which equipment, which process step does a value belong to? Are data definitions consistent across sites? Without that link, a value stays meaningless.

Dimension 4

Governance

Who owns which data, who is responsible for maintaining it? Governance is not a one-time setup but a living operating model — from the local pilot to the enterprise-wide rule.

What it costs to skip this step

Projects that start with “let's collect data first” become platform initiatives instead of value initiatives. The architecture gets delivered — but daily work does not change. Operators still keep their shift handover in Excel. Quality reports still take too long. Trends are still pulled by hand.

And if there is no way to turn the first use case into the next, the platform stagnates. Lessons are not shared, new ideas not picked up. In the end the expensive new data platform becomes the next legacy burden — the same trap, just with newer technology.

From experience: data quality problems can almost always be brought under control. Missing business context is far harder to retrofit later. That is why contextualisation belongs at the beginning — not at the end, once the platform is already standing.

The common denominator

So the problem is rarely the volume of data. It is the missing shared context, the unclear ownership and the missing common language between OT (Operational Technology, the technology of production) and IT (Information Technology, the enterprise IT).

And that is exactly why it takes two perspectives that rarely sit at the same table: the technical foundation — connectivity, OT/IT integration, contextualisation — and the organisational direction — the why, the governance, the adoption. Without either one, even the best data platform stays a technically sound but commercially inconsequential project.

The next posts in this series walk the path step by step: from GxP-compliant connectivity, through the question of why data projects fail, to the cases where connected data turned into real decisions.

You may see some of these points differently from your own practice — that is the point. We are less interested in who is right than in where the common language is missing in your organisation. That is exactly what we would like to talk about in Berlin.

Read on · Part 2 of the series

GxP-compliant and still usable: why compliance is not a data problem

How data can flow without touching validated systems — and why compliance is mainly a question of context.

Read the post →

We will be discussing this at Pharma MES Europe 2026

In Berlin, cts Group and Xenium AG show how fragmented production data becomes a usable foundation — from strategy to the running plant.

Book a meeting at the stand