Predictive maintenance is optimising for uptime. Your safety layer is not.
Condition monitoring finds degradation that announces itself. A proof test finds dangerous failures that do not, which is why PdM coverage cannot be traded for a longer proof test interval without changing the PFD you certified.
A disclosure before the argument, because it belongs in the open: the operator of this publication has affiliated business interests in instrumentation, predictive analytics, and plant digital models, the category this piece examines. That is precisely why the standard here is evidence and independence, not enthusiasm. If the argument below cuts against the category, so be it. The floor does not care who is selling.
Predictive maintenance has earned its position. Vibration, thermography, oil analysis, motor current signature, and now a great deal of pattern recognition on historian data have made it genuinely possible to see a bearing, a pump, or a heat exchanger degrading well before it fails. The business case is clean, the savings are measurable, and adoption keeps widening.
The pressure worth watching is what happens when that capability reaches the safety-instrumented layer, and someone reasonable asks the obvious question: if we can monitor the condition of this equipment continuously, why are we still shutting the unit down to proof test the trip.
It is a fair question. The answer is not “because the standard says so.” The answer is that the two activities are looking for different failures, and only one of them is looking for the failure that defeats the protection.
What a proof test is actually for
Functional safety for the process sector runs on IEC 61511, the sector-specific application of IEC 61508. In the United States, safety instrumented systems also sit inside the mechanical integrity and RAGAGEP obligations of the federal Process Safety Management standard, 29 CFR 1910.119. In the United Kingdom, HSE has published operational guidance specifically on proof testing of safety instrumented systems, which is worth reading directly because it deals with the practical failure patterns rather than the theory.
A safety instrumented function spends nearly all of its life doing nothing. It sits there, and the plant runs. Because it is dormant, a failure in it produces no symptom. The valve that will not close on demand closes nothing today, because nothing asked it to. That class of failure has a name in the standard, the dangerous undetected failure, and the proof test exists for one purpose: to find those before a real demand does.
This is why the arithmetic works the way it does. For a low-demand function, the simplified relationship is that average probability of failure on demand is approximately the dangerous undetected failure rate multiplied by the proof test interval, divided by two. PFDavg is roughly (λDU × TI) / 2. Read that as a physical statement rather than a formula: risk accumulates with the time since you last looked. Double the interval and you roughly double the probability that the function is already dead when it is needed. The SIL you certified is a claim about how often you look.
Why condition data does not substitute
Predictive maintenance is very good at a specific class of failure: progressive, mechanical degradation that emits a signal. Bearing wear, imbalance, cavitation, fouling, thermal drift. It is far weaker on the failure modes that actually matter in a dormant protective loop.
A final element that will stroke fine and still not seat under process conditions. A solenoid that has quietly seized after years of no movement. An impulse line that has plugged, so the transmitter reports a comfortable and completely fictional pressure. A logic solver output card that will not change state. A bypass or override left in place after a maintenance activity, or a force set during commissioning that nobody removed. An instrument correctly reading a variable it is no longer physically connected to in the way the design assumed.
Several of those produce no degradation signal at all, because nothing is degrading. The component is in a wrong state, not a worn one. And some of them are introduced by human action rather than time, which no trend line anticipates.
There is a second point that gets lost. A proof test tests the function end to end: sensor, logic, final element, and the trip actually happening. Condition monitoring watches components. A loop can be made of individually healthy components and still not trip, because the failure lives in the logic, the wiring, the bypass state, or the interaction. Proof test coverage is a defined and auditable property of a test procedure. It is not something a condition-monitoring deployment inherits by being nearby.
The version that is legitimate
None of this makes PdM irrelevant to the safety layer. It makes its role specific. Continuous diagnostics genuinely can convert some dangerous undetected failures into dangerous detected ones, and where that is real, the standard’s own framework accounts for it through diagnostic coverage. That is the honest route: demonstrate the coverage, put it in the calculation, and let the interval follow from the resulting PFDavg. Partial stroke testing on a valve is the familiar example of the same logic, and note that it is credited with partial coverage, not full.
The illegitimate route is the one that will show up in budget conversations: we now have analytics on this equipment, so we can safely stretch the turnaround interval. That is a change to the risk claim, argued from an efficiency benefit, usually by people who are being genuinely helpful and are optimising a different objective. Availability and protection are not the same target, and the gap between them is exactly where a dormant failure sits.
Before the interval moves
When anyone proposes extending a proof test interval on the strength of monitoring or analytics, ask: which specific dangerous undetected failure modes does this monitoring actually reveal, what proof test coverage are we claiming, and has the PFDavg been recalculated and re-verified against the target SIL with that claim in it? Require the answer in failure modes, not in platform capability. Then check the modes it cannot see: will it detect a final element that strokes but does not seat, a plugged impulse line, a seized solenoid, or a bypass left in. Finally ask who signed the change, and whether that person owns the functional safety case or the maintenance budget. If the analysis has not been redone and re-verified, the interval is being set by a business case, and the SIL on the drawing is no longer a claim anyone has checked.
Predictive maintenance is going to keep getting better and it belongs in the plant. Just keep the objectives separate. Uptime work asks when this will break. The safety layer asks whether the thing that has not moved in three years will move when it has to, and only a test that makes it move can answer that.