Counting observations isn't reducing incidents
If your BBS KPI is observation volume, you will get observations, not fewer injuries.
Behaviour-based safety (BBS) rests on a defensible idea: most incidents involve an at-risk act, so watch for the acts and coach them out. The trouble is what happens the moment you attach a number to it. Set a target of two observations per supervisor per week, and you have defined success as the production of observation cards. Cards will appear. Whether risk moved is a separate question the target does not ask.
The metric answers a question you did not mean to ask
An observation count is a proxy. It stands in for “we are watching the work and correcting hazards,” but it measures only “forms were completed.” Those come apart fast under pressure. When the KPI is volume, the cheapest way to hit it is to observe easy, low-stakes behaviours, a colleague already wearing safety glasses at a desk, and log the checkmark. The hard observations, near live energy or at height, are exactly the ones a volume target discourages, because they cost time and social capital per card.
There is a nastier failure mode regulators have named. When observation programs are paired with rate-based incentives, bonuses tied to lower reported injuries, pressure can flip from finding hazards to suppressing reports. OSHA’s 2012 interpretation on safety incentive policies warns that programs disqualifying workers from a reward after an injury “might have dissuaded reasonable workers from” reporting, and that vague rules such as requiring an employee to “maintain situational awareness” can be a pretext to blame the injured (US; anti-retaliation basis 29 CFR 1904.35). A program can post a rising observation curve and a falling recordable rate at once, and the second can be an artifact of the first.
The premise underneath is doing more work than it can carry
Set the metric aside for a moment, because there is a prior question about the model itself.
BBS starts from the observation that an at-risk act is present in most incidents. That is true and almost uninformative, because an at-risk act is also present in most of the thousands of task repetitions that ended fine. What separates the two is usually not the act. It is whether a condition was there to catch it: the guard, the isolation, the barrier, the layout that made the safe way the easy way.
Which places behaviour coaching where the NIOSH hierarchy of controls has always placed it, in the administrative tier, near the bottom, dependent on a person choosing correctly under pressure every time. That is not an argument for abandoning it. It is an argument against letting it become the visible bulk of the safety programme, because a site can run a very busy observation programme while the engineering work that would remove the choice entirely goes unfunded.
The tell is what your observations produce. If most cards end in a conversation, the programme is operating as an administrative control. If a meaningful share end in a physical change to the workplace, it is functioning as a hazard-finding instrument, which is a genuinely different and more valuable thing.
What a leading indicator has to do to earn the name
The word “leading” gets applied to any number that is not an injury count, which is how observation volume ended up on executive dashboards. A real leading indicator has to clear three bars, and volume clears one.
It must be measurable before the outcome, which observation counts are. It must be mechanistically linked to the outcome, meaning there is a chain you can describe from the number moving to the injury not happening. And it must be actionable by the people being measured, so that improving it requires improving safety rather than improving reporting.
Observation volume passes the first and fails the other two. Hazards corrected and verified passes all three: the chain is direct, the supervisor can affect it only by finding and fixing real things, and it cannot be inflated by watching someone type safely.
Before you buy
Pull last quarter's observation records and ask one thing: how many produced a hazard correction that outlived the conversation, a guard added, a procedure changed, a tool retired? If that number is near zero while the count is high, you are measuring compliance with a quota, not risk.
Count the fixes, not the forms
None of this condemns watching the work. It condemns paying for it by the pound. A leading indicator is only leading if it moves before the outcome and points at something you then fix. Count the fixes, not the forms.
The redesign is not complicated. Drop the per-supervisor card quota entirely. Track corrections closed and verified, with the verification done by someone other than the person who logged it. Report the median time from observation to correction, because a hazard found and left open for six weeks was not really found. And keep the whole thing structurally separate from any injury-rate bonus, so that nobody in the system has a financial reason for a report not to exist.
Then the number on the dashboard is one you would be happy to have audited, which is the only test of a safety metric that has ever mattered.