Skip to content
StrataHub

Manufacturing · March 23, 2026 · 5 min read

Making Machine Data Useful for Manufacturing AI

Factories generate terabytes of sensor and PLC data that mostly goes nowhere. The gap between 'we have the data' and 'a model can learn from it' is a data engineering problem with a known shape — here is how we close it.

Every factory tour includes the same sentence: "We collect all the data." And it's true — the historian has been logging tags for a decade, the PLCs emit thousands of signals, the newer machines speak OPC UA. Then we ask to see the data behind last quarter's scrap spike, and it takes three people two weeks to assemble a CSV nobody fully trusts.

Having data and having usable data are separated by a canyon. Every manufacturing AI use case we build — predictive maintenance, quality prediction, energy optimization, the scheduling agents we've written about before — dies or thrives on whether that canyon gets bridged. So let's talk about the bridge.

Why machine data is hostile by default

Machine data isn't dirty the way CRM data is dirty. It has its own failure modes.

No shared clock. The PLC timestamps in local time, the MES in UTC, the vision system in whatever its vendor chose, and the historian interpolates on retrieval. A model correlating spindle vibration with a defect logged "at the same time" can be off by minutes — which is forever at line speed.

Tags without meaning. PLC7_DB12.DBW44 is not a feature. Which machine, which sensor, which unit, what does 4,095 mean when the sensor saturates? That context lives in a controls engineer's head or an Excel export from 2017.

Context-free streams. The sensor stream doesn't know which product was running, which recipe, which operator, which batch of raw material. Without joining machine data to production context, you can detect that something changed — never why.

Compression that lies. Historians deadband and swing-door compress data to save storage. Fine for trending on a dashboard; fatal for a model hunting for high-frequency signatures that were thrown away at ingest.

The contextualization layer is the product

The fix is not "put it all in a data lake." A lake full of uncontextualized tag data is the same problem with a bigger bill. The asset that makes machine data useful is a contextualization layer — and building it is the real project.

Concretely, that means four things.

An asset model. A structured registry — ISA-95 hierarchy or a pragmatic subset — mapping every tag to an asset, a sensor type, engineering units, valid ranges, and failure behavior. We build this incrementally, use-case first: model the three machines the first project needs, not all 400.

Time alignment as a service. One convention (UTC, always), documented sensor latencies, and explicit resampling policies per signal class. High-frequency vibration gets windowed feature extraction near the source; slow-moving temperatures get honest interpolation rules.

Production-context joins. Every window of machine data joinable to work order, product, recipe, shift, and material lot. This is usually a MES/ERP integration problem, and it is where genealogy questions ("which coils fed the parts that failed?") become answerable.

Quality gates at ingest. Stuck-at detection, range checks, gap accounting, unit-drift alarms. A silent sensor failure that feeds flatlined values into a predictive-maintenance model for three weeks will do more damage to trust than any model bug.

Rule of thumb from our engagements: for a first manufacturing-AI use case, 60–70% of the effort is this layer. The good news is it amortizes — the second use case on the same lines typically ships in a third of the time.

Architecture that has worked

We keep the pattern boring and repeatable. Edge collection via OPC UA or MQTT (Sparkplug B where the estate supports it), buffered locally so network blips don't create gaps. A streaming backbone into a time-series-competent store — TimescaleDB, or a lakehouse with proper partitioning at larger scale. The contextualization layer materialized as governed, versioned tables: asset-hours, joined and cleaned, that every downstream model consumes. Nobody trains against raw tags.

Two production disciplines matter as much as the stack. Lineage: when a model flags an anomaly, an engineer will ask "show me the raw signal" — and the path from prediction back through features to source tags must be traversable in minutes. Drift monitoring on inputs, not just outputs: sensors get recalibrated, PLC programs get patched, a maintenance tech swaps a transducer with a different response curve. Each of these silently shifts distributions. We monitor per-signal statistics and alert on shifts before the models degrade, because in a factory the data changes for physical reasons on physical schedules.

Start narrow, prove the loop

The failure mode we see most is the two-year "unified namespace" program that connects everything and improves nothing. The version that works inverts it: pick one line and one question with money attached — why does changeover scrap vary 3x across shifts? which spindle failures give us warning? — and build the full path for just that: edge to context layer to model to a screen someone acts on.

That scope fits a 4–6 week Pilot: real tags, real joins, a real answer, and the first slice of a contextualization layer that every subsequent use case inherits. Co-Build then extends the asset model line by line, use case by use case, with the data quality gates and observability going in as standard equipment rather than retrofits.

The historian already holds ten years of data. The question is whether year eleven finally gets to do some work.

Work with us

Shipping something like this?

We co-build production AI systems with enterprise teams — pilots in 4-6 weeks.