Independent · Judgment-led Reference publication · Industrial safety Follow · 4,222
From the Floor.

Ground truth for safe work.

The Scout

The Headset Logs Hours. The Floor Logs Behavior. Only One of Those Is Training.

VR safety training's value is established only by measured behavior change on the floor. Completion rates, engagement scores, and headset hours cannot stand in for it, no matter how good the simulation feels.

August 3, 2026

A four-rung Kirkpatrick ladder with L1 Reaction and L2 Learning filled as headset metrics and L3 Behavior and L4 Results outlined as the safety value the headset does not measure.

A worker puts on a headset, walks through a virtual lockout sequence, feels the small jolt of a simulated arc flash, and takes the headset off visibly affected. Everyone in the room agrees something happened. The engagement was real. The question that decides whether it mattered is one nobody in that room can answer yet: six weeks from now, standing at an actual disconnect, does that worker do the thing differently. That is the only measurement that counts, and it is the one the technology does not produce for you.

A disclosure before the argument, because it belongs in the open: the operator of this publication is active in the VR safety training market. That is precisely why the standard here is evidence and independence, not enthusiasm. If the argument below cuts against a category the operator sells, so be it. The floor does not care who is selling.

Immersion is a real advantage, at the levels where it operates

Let us give VR its due, because the case for it is genuine and it is easy to caricature. Immersive simulation is good at attention. It is good at letting people practice a hazardous sequence without the hazard. It is good at completion, because people finish things they find engaging, and it is good at consistency, because every learner gets the same scenario rather than whatever the instructor remembered to cover that morning. These are not small things. A training that people actually finish, and that puts hands on a procedure before the procedure can hurt them, is a better starting point than a slideshow and a signature.

But notice what all of those advantages have in common. They are about the experience of the training and the knowledge it delivers. They are not, yet, about what happens on the floor.

The framework that has always sorted this out

The distinction is not new and it is not controversial. The most widely used framework in training evaluation, the Kirkpatrick model, has sorted training outcomes into four levels for more than sixty years. Level 1 is Reaction: did learners find it engaging. Level 2 is Learning: did they acquire the knowledge. Level 3 is Behavior: are they actually doing the work differently back on the job. Level 4 is Results: did the outcomes that matter, including incidents, actually move.

The whole discipline of the model is in one uncomfortable observation: each level is harder to measure than the one before, and each is more meaningful than the one before, and organizations reliably stop measuring right at the point where it starts to matter. VR is spectacularly good at Level 1. It is often good at Level 2. Almost every metric a headset generates natively, completion, time in scenario, quiz scores, self-reported confidence, engagement, lives at Levels 1 and 2. And the safety value of a training does not live there. It lives at Level 3, behavior on the floor, and Level 4, the incidents that did or did not happen.

This is the falsifiable point: VR safety training’s value is established only by measured on-floor behavior change, and completion or engagement metrics cannot stand in for it. If a program can show behavior change at the point of work and cannot show it was the training that moved it, the value is asserted, not demonstrated. If it can only show headset hours and satisfaction scores, it has not even reached the level where safety value is defined.

Diagnostic

Look at the dashboard your VR program actually reports against, and label each metric with its Kirkpatrick level. Completion rate, engagement score, hours logged, learner satisfaction: those are Level 1 and Level 2, and they are the easy ones to collect, which is exactly why they crowd out the rest. Now ask the question the vendor slide will not: what is your Level 3 measure, the observed behavior on the floor that this training was supposed to change, and how would you know if it did not change at all? If you cannot name that measure, you are buying a training experience and calling it a safety outcome.

Why the substitution is tempting, and why it is a trap

The reason completion metrics stand in for behavior is not that anyone is being dishonest. It is that Level 1 and Level 2 data is cheap, immediate, and flattering, and Level 3 data is expensive, delayed, and sometimes disappointing. A completion rate arrives the moment the session ends. A behavior change has to be observed at the point of work, weeks later, by someone watching whether the procedure is actually followed, and then it has to be connected back to the training rather than to a supervisor’s reminder or a recent near miss.

That is real work. It looks like structured behavioral observation, or task-based competency checks done at the job, or leading indicators tied to the specific behavior the module targeted, measured before and after with enough discipline to tell training effect apart from everything else changing on a busy floor. None of it is generated by the headset. All of it is generated by the organization deciding the question is worth answering.

This is where “compliance is not safety” earns its keep. A completed VR module produces a clean training record. A regulator, in most jurisdictions including the United States, will accept documented, completed training as evidence the requirement was met. That record is a compliance artifact. It is not evidence of behavior change, and treating the two as equivalent is the exact substitution the Kirkpatrick model was built to prevent. The audit passes at Level 1. The worker at the disconnect is a Level 3 question, and the disconnect does not read your training log.

What “what works” would require

The honest version is not anti-VR. It is VR held to the standard any training method should meet. Define the specific on-floor behavior the module is supposed to change. Measure that behavior before you deploy. Deploy. Measure again, at the job, with enough rigor to attribute the change. If the behavior moves and holds, you have earned the claim, and immersion may well be why it moved. If it does not, the engagement was real and the transfer was not, and no completion rate should be allowed to hide the difference.

Immersion is a promising means. On-floor behavior is the end. Any program, from any operator, that reports the first and stays quiet about the second has told you which one it could measure, not which one worked.