AUTONOMOUS INDUSTRIAL DIGITAL TWIN
The Business Problem
Where the data is trapped
A manufacturing site already generates more data than it uses. The controllers know the state of every
valve and every drive. The historian holds years of process values. The manufacturing execution system
knows what was made and when. Quality results sit in the laboratory system, maintenance in the
computerised maintenance management system, cost and planning in the enterprise resource planning
system.
Each of those systems is correct on its own and none of them can answer the question that matters. When
a batch fails final inspection, nobody can say which process window produced it, which machine was
drifting at the time, and whether the same conditions are present on the line right now.
What the site was doing instead
- Export to a spreadsheet. A process engineer pulls a historian trend, joins it by hand to a quality result
and reaches a conclusion that cannot be repeated next month. - Run to failure or run to calendar. Maintenance is done on a fixed interval because condition data exists
but has never been connected to failure outcomes. - Inspect at the end. Defects are found after the value has already been added, so scrap is discovered at
the most expensive possible moment. - Optimise one line at a time. Because nobody has a view that spans lines, sites and the supply that feeds
them.
The result is a site that is heavily instrumented and still managed by experience.
Where the cost sits
| Cost | Driver |
|---|---|
| Unplanned downtime | Failures that had a detectable signature weeks earlier in data nobody was reading. |
| Scrap and rework | Process drift detected at inspection instead of at the point it started. |
| Over maintenance | Components replaced on a calendar while healthy, and others missed while degrading. |
| Yield loss | Lines run at a conservative setpoint because nobody can prove the wider window is safe. |
| Energy per unit | Utilities treated as a fixed overhead rather than as a controllable input. |
| Changeover time | Sequencing decided by habit rather than by the actual changeover matrix. |
| Knowledge loss | Failure mode experience leaving the site with retirement. |
Solution Overview
Data is collected at the edge, contextualised against one asset hierarchy, curated into a lakehouse, and used to maintain a twin state for every asset that matters. Models run against the twin rather than against
raw tags, which is what makes them portable between lines and between sites instead of being rebuilt for each one
Layered architecture
The advisory boundary
The platform reads the plant continuously and writes to it only through a path the control engineers own. That boundary is what makes the system approvable by the operations technology security team and by the process safety authority.
| The platform does | The platform never does |
|---|---|
| Read tags, batches and events | Write to a safety instrumented system |
| Predict a failure and rank it | Stop a line on its own authority |
| Recommend a setpoint change | Move a setpoint outside the approved envelope |
| Propose a production sequence | Release a sequence without planner approval |
| Raise a quality alarm | Release or reject a batch |
| Prepare a work order | Bypass a permit to work |
Edge Acquisition and Protocols
Protocol coverage
| Protocol | Typical source |
|---|---|
| OPC UA | Modern controllers and the plant's structured data layer, including the type information that makes contextualisation possible. |
| OPC DA and classic | Older cells where a bridge is needed rather than a rewrite. |
| Modbus TCP and RTU | Drives, instruments, protection relays and legacy equipment. |
| Siemens S7 and Rockwell | Controller families where a native driver is faster and lighter than a gateway. |
| MQTT with Sparkplug | Existing telemetry brokers, and the preferred shape for new edge devices because state and birth certificates come with it. |
| Historian query | Bulk history for model training and for trend context that the live tags cannot give. |
| MES and ERP | Production orders, materials, batches and cost, through APIs or change capture. |
| Vision and audio | Cameras and acoustic sensors, processed at the edge with only the result and the exception image sent upstream. |
Context and the Asset Model
A raw tag means nothing on its own
A value of 71 degrees is not information. The same value becomes information when the platform knows it came from the drive end bearing of pump six, that pump six is part of the cooling loop on line three, that line three was running product B on order 44817 at the time, that the ambient was 38 degrees, and that
this reading has risen from 58 over six weeks.
The asset model is what carries all of that. It follows the standard equipment hierarchy of enterprise, site, area, line, unit and equipment, and every signal is bound to a node in it. Models are then written against the node type rather than against a tag list, which is what allows a pump model built on line three to be
deployed on line seven without rewriting it.
What context is joined
Material genealogy Maintenance history Quality result Ambient conditions
Energy meter
The contextualised state object
The Digital Twin Engine
What the twin actually is
The twin is three things held together. A current state assembled from the plant, a model of how the asset should behave under those conditions, and the residual between the two. The residual is where all the value sits, because a machine that is degrading looks normal against a fixed threshold and abnormal against its own expected behaviour.
Keeping the twin honest
A twin that is not corrected drifts away from the plant within weeks. Parameters are recalibrated on a schedule against recent normal operation, and every recalibration is versioned, so a change in the twin can be distinguished from a change in the machine.
Fidelity is chosen per asset, not per plan
| Tier | Where it is worth it |
|---|---|
| Statistical | Large populations of similar low criticality assets. Cheap, and enough to rank them. |
| Behavioural | Assets with a known operating curve. Expected value derived from load and conditions, residual tracked. |
| Physics based | Critical or expensive equipment where a mass and energy balance is worth building and can be validated. |
| Line level | A discrete event model of the line, used for throughput, buffer and changeover questions rather than for equipment health. |
| Network level | Supply, inventory and multi site allocation, run as scenarios rather than continuously. |
The Model Estate
What runs, and what it is for
| Model | Job and shape |
|---|---|
| Predictive maintenance | Remaining useful life and failure probability per asset class, trained on residuals and on the maintenance record rather than on raw tags, so it transfers between sites. |
| Anomaly detection | Multivariate detection on the twin residual, which catches combinations that no single tag threshold would raise. |
| Visual inspection | Defect detection and classification at the line, running at the edge with only results and exception images sent upstream. |
| Quality prediction | Predicting the laboratory result from process conditions, so a drift is visible hours before the sample is taken. |
| Root cause analysis | Ranking the process variables most associated with a quality excursion, with the correlation stated as correlation. |
| Energy and utilities | Forecasting demand and identifying the load that can be shifted without touching output. |
| Throughput and scheduling | Sequencing against real changeover times, real availability and real yield rather than against the planning assumptions. |
| Supply and inventory | Demand sensing and buffer sizing across the network. |
Training data is the constraint
Every predictive maintenance programme meets the same wall. Failures are rare, which is the point of the plant, and a model needs examples. Three things make it workable.
- Label from the maintenance record, carefully. A work order tells you something failed. It rarely tells you
when degradation started, so the labelling window is agreed with the reliability engineers rather than
assumed. - Pool across identical assets. Forty pumps of the same type on one site give a usable population where
one pump gives nothing. - Use the residual, not the raw signal. Because the residual is already normalised for duty and conditions,
which is most of the variance that would otherwise swamp the failure signal.
Lifecycle
Agents and the Reasoning Layer
Bounded copilots, not a general assistant
Each agent has a defined question set, a defined set of tools and a defined authority limit. It reads the twin, searches the plant document corpus, queries the lakehouse and produces an answer with references that the engineer can open.
| Agent | Scope and limit |
|---|---|
|
Maintenance copilot |
Diagnoses a symptom against failure modes and asset history, drafts the work order with parts and duration. It cannot release the work order. |
|
Operations copilot |
Explains why a line is behind, which constraint is binding and what recovering it is worth. It cannot change the plan. |
| Quality copilot | Traces an excursion back through process conditions and material genealogy. It cannot release or reject a batch. |
|
Scheduling copilot |
Runs sequence scenarios against the twin and presents the trade off. The planner chooses. |
| Shift handover | Writes the handover from events, alarms and interventions, for the outgoing supervisor to correct and sign. |
| Energy copilot | Identifies shiftable load and states the effect on output before it states the saving. |
The plant corpus
Failure mode analyses Vendor bulletins Change records Shift logs
Grounding contract
| condition | Required behaviour |
|---|---|
| Supported by documentation | Diagnoses a symptom against failure modes and asset history, drafts the work order with parts and duration. It cannot release the work order. |
|
Operations copilot |
Explains why a line is behind, which constraint is binding and what recovering it is worth. It cannot change the plan. |
| Quality copilot | Traces an excursion back through process conditions and material genealogy. It cannot release or reject a batch. |
|
Scheduling copilot |
Runs sequence scenarios against the twin and presents the trade off. The planner chooses. |
| Shift handover | Writes the handover from events, alarms and interventions, for the outgoing supervisor to correct and sign. |
| Energy copilot | Identifies shiftable load and states the effect on output before it states the saving. |
Worked Example, End to End
Security and OT to IT Convergence
Network position
The platform’s plant footprint is an operations technology asset and it is treated as one. It sits where a monitoring device sits, not where a control device sits, and the flow of data is outward.
| Control | Implementation |
|---|---|
| Zone and conduit | Collectors placed in a defined zone with declared conduits, rather than introducing a new trust boundary across the site. |
| Read only clients | No write function compiled into the acquisition layer for any signal outside an explicitly approved write list. |
|
Unidirectional option |
A data diode where the site requires physical enforcement rather than configuration. |
| Outbound only | The plant initiates the connection to the cloud. No inbound path from the enterprise into the control network. |
| Authentication | Certificate based, with OPC UA sessions signed and encrypted rather than left on the anonymous endpoint. |
| Patching | Edge software updated through the site's existing controlled change process, on the site's maintenance window. |
|
Segregation of duties |
The team that can change a control system is not the team that can change the platform. |
Why this is the whole approval argument
An operations technology security team assesses a new system on two questions. Can it change anything in the plant, and can anything reach the plant through it. This architecture answers no to both by construction rather than by configuration, which is what makes the assessment tractable and short
Governance on the enterprise side
- One catalogue across sites, with a site seeing its own data by row policy and the group seeing the
aggregate, so multi site analytics does not require copying data around. - Recipes, formulations and process windows treated as sensitive, masked by default and released by role, because this is usually the most commercially valuable data the company owns.
- Lineage from a board level number back to a tag, which is what makes an operations claim auditable.
- Retention by data class, since high rate vibration data and monthly cost data do not need the same
treatment or the same cost.
How We Delivered It
The phases we ran
| Phase | Scope | What it produced |
|---|---|---|
|
Phase 0 Assessment |
Asset and tag survey, protocol survey, historian audit, network placement, security engagement, one asset class chosen. | Tag to asset mapping quality measured and network design agreed. |
|
Phase 1 Foundation |
Edge collectors, streaming into bronze, asset hierarchy built, silver and gold for one line. | Data reconciled against the historian and the execution system for a full month. |
|
Phase 2 Twin |
Behaviour models and residuals for the chosen asset class, calibrated on recent normal operation. | Residual stable on healthy assets and elevated on assets known to be degrading. |
|
Phase 3 Advisory pilot |
Predictive maintenance and quality prediction in shadow, with a small reliability and process group reviewing every output. | Precision and lead time accepted by the reliability engineers. |
|
Phase 4 Line rollout |
Additional asset classes and lines, copilots released, command centre live. | Downtime and scrap targets held for a quarter. |
|
Phase 5 Multi site |
Model and pipeline packaging standardised so a new site is a deployment rather than a project. | Repeatable site onboarding package. |
How it runs now
Benefits and Measurement
Benefits realised
KPI framework
| Measure | What it tells you |
|---|---|
|
Overall equipment effectiveness |
The headline number, split into availability, performance and quality so the movement can be attributed. |
|
Unplanned downtime hours |
The outcome the maintenance case rests on. |
| Prediction lead time | How much warning the site actually got, which decides whether the warning was useful. |
| Prediction precision | Share of alerts confirmed when the machine was opened. Low precision destroys trust faster than low recall. |
| Scrap and rework rate | Quality outcome, measured per line and per product. |
| First pass yield | Whether the process is being held inside the window rather than corrected after. |
| Energy per unit produced | Utilities normalised for output, which is the only comparable form. |
| Tag to asset coverage | How much of the plant the platform can actually reason about. |
Scale and Performance
Scale the platform was built to carry
| Metric | Enterprise scale |
|---|---|
| Manufacturing sites | 20 to 100 |
| Connected devices | 250,000 to 1 million |
| Sensor readings | 150,000 to 500,000 per second |
| Daily data ingestion | 50 to 150 TB |
| Video streams | 1,000 to 4,000 |
| Robots connected | 500 to 5,000 |
| Digital twins maintained | 20,000 to 100,000 |
Engineering targets
| Measure | Target |
|---|---|
| Streaming latency | Under 2 seconds |
| Predictive maintenance inference | Under 1 second |
| Computer vision processing | 25 to 35 frames per second |
| Data availability | 99.9 percent |
| Model deployment frequency | Weekly |
| Equipment data synchronisation | Near real time |
Business improvement ranges
| Measure | Expected improvement |
|---|---|
| Unplanned downtime | 20 to 35 percent reduction |
| Maintenance cost | 15 to 25 percent reduction |
| Production throughput | 10 to 18 percent increase |
| Overall equipment effectiveness | 8 to 15 point improvement |
| Product defect rate | 18 to 30 percent reduction |
| Energy consumption | 10 to 18 percent reduction |
| Spare parts inventory | 15 to 20 percent reduction |
| Maintenance planning time | 40 to 60 percent faster |
Platform Capability Across the Estate
| Platform metric | Typical target |
|---|---|
| Data ingestion | 20 to 150 TB per day |
| Structured streaming throughput | 20,000 to 500,000 events per second |
| Historical lakehouse | 2 to 30 PB |
| Delta tables | 5,000 to 25,000 |
| Models under management | 100 to 500 |
| Feature store features | 10,000 to 100,000 |
| Vector embeddings | 100 million to 2 billion |
| Platform metric | Typical target |
|---|---|
| Daily inference requests | 10 to 100 million |
| Enterprise users | 5,000 to 50,000 |
| Platform availability | 99.9 to 99.95 percent |
| Automated data quality checks | Above 95 percent of published tables |
| Governance coverage | 100 percent of production datasets |
| Mean time to detect a data issue | Under 15 minutes |
What We Learned
The hard problems, and what we did about them
| Problem | What we did |
|---|---|
| Tag to asset mapping | Surveyed and measured in Phase 0 with an explicit budget. This is the dominant driver of first phase duration and it is knowable before anything is committed. |
| Historian compression | Deadband and compression settings recorded per signal, and training data drawn the same way the live path will see it. |
| Few failure examples | Pool across identical assets, label with the reliability engineers, and start with anomaly detection where supervised learning is not yet possible. |
| Plant network constraints | Edge buffering sized for the longest realistic outage, and bandwidth calculated from the real sampling plan rather than assumed. |
| Operations technology approval | Security team engaged in Phase 0, not at deployment, with read only enforced in the client and in the network. |
| Operator trust | Advisory pilot with a small group, mandatory citations, and every prediction closed out against what was found. |
| Model decay after plant changes | Change records fed into the platform so a modification invalidates the affected twin rather than silently degrading it. |
| Cost of high rate data | Feature extraction at the edge for vibration and acoustic signals, with raw waveform kept only around events. |
What we settled before writing any code
- The first asset class. Which line and which asset class would go first, chosen so the value was
measurable inside two quarters. - Control network exposure. What was available, through which protocol, and whether an OPC UA server
already existed. - The historian. Which product, what retention, and whether it could be queried in bulk without affecting
operations. - The asset hierarchy. Whether a maintained hierarchy existed anywhere, and who owned it.
- Maintenance history. Whether the records were detailed enough to label failures, and who could
authorise their use. - Security ownership. Who owned operations technology approval, and what their assessment process
actually was. - Closed loop intent. Whether closed loop control was an eventual objective, and whether the control
system supported a bounded envelope. - Site count. How many sites were in scope, and how similar their control systems were.