Independent · Judgment-led Reference publication · Industrial safety Follow · 4,222
From the Floor.

Ground truth for safe work.

The Scout

Wearables promised fewer slips. What the evidence actually shows.

A 12-week pilot can't tell you whether a sensor works, because the attention around the pilot moves the numbers as much as the device does.

June 26, 2026 · updated August 8, 2026

Two bars showing a wearable pilot group and a control group that got equal attention but no device reaching similar outcomes, isolating the small device effect.

Sensor wearables, posture monitors, proximity tags, fatigue trackers, are among the fastest-moving categories in safety tech. The pitch is intuitive: measure the risky movement, nudge the worker, prevent the injury. The demo comes with a pilot result: incidents down, posture scores up, everyone impressed.

The harder question is whether that result is the device working, or the well-documented effect of any program in which people know they’re being watched and supported.

The pilot problem

NIOSH and others have flagged the difficulty of evaluating emerging sensor technologies, and the evaluation problem is well understood: short pilots, motivated volunteers, and concentrated management attention all push outcomes in the same favourable direction, independent of the hardware. A twelve-week trial with the safety team hovering is close to the ideal conditions for any intervention to look good. Strap on a placebo and you might see the numbers move too.

That doesn’t mean wearables don’t work. It means a pilot designed to sell one can’t tell you whether they do. The measured effect you care about is the device’s marginal contribution, what it adds beyond the attention, coaching, and Hawthorne glow that came free with the pilot.

Why the pilot started is part of the result

There is a second effect underneath the first, and it is purely statistical.

Sites do not commission wearable pilots at random. They commission them after a bad run: a cluster of manual-handling claims, a serious near-miss, a spike that got attention upstairs. The pilot begins precisely at the point where the numbers were unusually high.

Unusually high numbers tend to come back down on their own. That is regression to the mean, and on a site with a small number of recordable events it is a large effect, not a footnote. The same arithmetic that makes a low incident rate twitchy makes a high one temporary. Any intervention introduced at the peak inherits the recovery and gets credited with it.

So the honest baseline is not the twelve weeks before the pilot. It is the site’s own multi-year run rate, and a comparable area that had the same bad quarter and received nothing.

Twelve weeks cannot detect what it claims to measure

The third problem is one of arithmetic rather than psychology, and it is the one that rarely gets raised in the room.

If a site records a handful of relevant injuries a year, a twelve-week window contains a fraction of one expected event. A study cannot demonstrate a reduction in something that was not going to happen very often during the observation period anyway. Whatever the pilot shows over that window, the count is dominated by chance rather than by the device, and no amount of dashboard polish changes that.

This is why serious evaluations of sensor technology measure exposure rather than outcome. Time spent in a high-risk trunk angle, number of lifts above a threshold, minutes inside a proximity envelope. Those are frequent enough to move detectably in twelve weeks, they are mechanistically linked to the injury you are trying to prevent, and they are what the device is actually instrumented to see. Asking a wearable pilot to prove an injury reduction is asking the wrong question of the right tool.

Before you scale a pilot

Ask the vendor for the control condition: what happened to a comparable crew that received the same attention and coaching but no device? If there wasn't one, you measured a program, not a product, and a program is far cheaper to run than a per-worker hardware lease.

The number that decays after the purchase order

One more figure is worth demanding, and vendors rarely volunteer it: wear time, plotted over the full pilot rather than averaged across it.

Adherence to any wearable is highest in week one, when the device is novel and the safety team is visible, and it falls from there. A pilot that reports 91% average wear time may be describing 99% in week two and something much worse by week twelve, and the trajectory tells you far more about life after rollout than the average does. Ask for the weekly curve. Ask what happened to the crews who quietly stopped charging the units.

The technology may well earn its place; sensing is genuinely improving. But the burden of proof sits with the isolated device effect, measured against a fair comparison, not with a pilot engineered to flatter it. Buy the evidence, not the enthusiasm.