Independent · Judgment-led Reference publication · Industrial safety Follow · 4,222
From the Floor.

Ground truth for safe work.

The Skeptic

The risk matrix is not a measurement

Strip the scores off ten hazards and have two competent people rescore them alone. If more than a third land on opposite sides of your action threshold, the matrix is not deciding anything.

September 16, 2026

A five by five risk matrix with a stepped action threshold line, showing two scorers placing the same hazard at the same severity but one cell apart on likelihood, landing on opposite sides of the line that decides whether a control is bought.

A residual risk has been open on the register for fourteen months. Engineering wants a fixed guard and an interlock. Operations wants a written procedure and a toolbox talk. The argument ends the moment someone opens the assessment file: likelihood 2, severity 4, product 8, amber, tolerable with existing controls. Nobody rescores it. The meeting moves on.

That is the whole problem in one scene. A grid that cannot measure anything just ended an argument it was never equipped to settle. And the file now records an 8 where the reasoning should be.

Ordinal labels do not multiply

The 5x5 grid asks two people to convert a judgement into a category, then treats the categories as if they were quantities. “Likely” and “moderate” are not defined amounts. They are words that different people attach to different ranges, and the ranges move with the last incident anyone remembers.

This is not a matter of taste. Louis Anthony Cox Jr set out the mathematics in What’s wrong with risk matrices?, published in Risk Analysis, volume 28, issue 2, April 2008, pages 497 to 512. The paper examines the properties of the matrix itself rather than any one organisation’s use of it, and the findings are specific.

Poor resolution: a typical matrix can correctly and unambiguously compare only a small fraction of randomly selected pairs of hazards, which the paper puts at under ten per cent. Range compression: quantitatively very different risks get identical ratings, which is why half your register is amber. Errors: the matrix can assign a higher qualitative rating to a quantitatively smaller risk. Where frequency and severity are negatively correlated, the paper describes matrices as “worse than useless”, producing decisions worse than random. And on inputs and outputs, the finding that should bother every EHS director: different users can obtain opposite ratings of the same quantitative risk.

The paper’s conclusion is not that matrices should be abolished. It is that they should be used with caution and with careful explanation of the judgements embedded in them. That is a narrower and more useful claim than either camp usually makes.

What the standards actually require, and where

Start with what people cite. The risk assessment technique catalogue most EHS teams refer to as “ISO 31010” is published as IEC 31010:2019, second edition, June 2019, replacing the 2009 first edition. Its stated purpose is guidance on the selection and application of techniques for assessing risk, with summaries of a range of them and references out to fuller descriptions. The 2019 revision added detail on planning, implementing, verifying and validating the use of a technique, and increased the number and range of techniques covered. That is the architecture of a catalogue. A catalogue does not nominate a winner, and a technique you have never verified or validated is not evidence of anything.

For machinery, ISO 12100:2010 sets out basic terminology, principles and a methodology for safety in machine design. Its published scope describes procedures for identifying hazards and for estimating and evaluating risks across the machine life cycle, and it gives guidance on documenting and verifying the risk assessment and risk reduction process. Nothing in that scope hands you a five by five grid. The standard asks you to show your work. A cell reference is not work shown.

Now the jurisdiction that decides whether you get cited. In the United States, general industry, 29 CFR 1910.132(d) requires the employer to assess the workplace for hazards that necessitate PPE, to select and have each affected employee use suitable PPE, to communicate the selection decisions, to select PPE that properly fits, and to verify the assessment through a written certification naming the workplace evaluated, the person certifying, the dates, and identifying the document as a certification of hazard assessment. Four required contents. No scoring scale, no grid, no threshold. Appendix B to subpart I is explicitly non-mandatory.

Process safety is the closest US OSHA comes to prescribing method. 29 CFR 1910.119(e) requires employers to use one or more appropriate methodologies from a named list: what-if, checklist, what-if/checklist, hazard and operability study, failure mode and effects analysis, fault tree analysis, or an appropriate equivalent. A risk matrix is not on that list. It is what teams often use to rank what those methods surface, which is a different job.

So the compliance position, in the United States, is that you must assess, and in places you must certify or document. The grid is a convention, not a requirement. Compliance is not safety, and here compliance does not even ask for the artefact most people think it demands.

The diagnostic

Open the last five residual risks your organisation formally accepted. For each one, read only what is recorded in the file, not what you remember from the room. Can a competent stranger reconstruct why the residual risk was accepted without asking anyone who was there? If yes, the matrix is doing what it is good at, sorting, and the reasoning lives somewhere else. If no, and the record is a score and a colour, then your acceptance decision has no stated basis, and the next person to inherit it will either re-litigate it from scratch or, more likely, leave it alone because it is already amber.

Use it to sort, not to justify

A matrix is a communication and triage device. It is genuinely good at that. Forty open items and a finite maintenance window is a sorting problem, and a shared grid gets a cross-functional group to a defensible running order faster than an argument does. Keep it for that.

Three limits make the difference between a triage tool and a liability.

It is not the justification for accepting a residual risk. The justification is the control reasoning: which options were considered against the hierarchy of controls as NIOSH sets it out, why elimination, substitution and engineering control were rejected or deferred, what evidence supports the adequacy of what was chosen, who accepted the remainder, and when it gets looked at again. Write that down. If the only record is a cell, you have recorded the output and deleted the thinking.

It is not a comparator across categories. Cox’s resolution finding bites hardest when you put a chemical exposure, a fall hazard and a machine guarding gap in the same ranked list and act as though the ordering means something. It orders your queue. It does not tell you which one kills someone first.

It is not a scale you should do arithmetic on. Multiplying two ordinal labels produces a number that looks like a quantity and is not one. If you must combine them, combine them as a lookup that you have defined and can defend, not as a product you treat as continuous.

The test

Take ten hazards off your live register. Strip the existing scores. Give the descriptions to two competent assessors who did not write them, and have them score independently, without conferring.

Count two things. Cell level disagreement, where the pairs land in different squares. And threshold level disagreement, where they land on opposite sides of the line that decides whether money is spent. Cell level noise is tolerable and expected. Threshold level noise is not.

If more than a third of the ten cross the action threshold depending on who is holding the pen, the matrix is not making your decisions. The scorer is. That has a concrete consequence at your next serious incident: the assessment file will be requested, it will show a score and a colour, and there will be nothing in it that explains why anyone believed the remaining risk was acceptable.