When the assistant answers to the regulation, not your procedure
Put 20 questions to your site assistant where your own written procedure is stricter than the rule. The share it answers to the regulation is your deployment decision.
Consider the exchange these deployments are built for. A technician on night shift needs to break into a pump line. The permit desk is in another building. On the wall is a tablet with a search box that reads documents and answers in plain sentences. He types: what is the isolation requirement here. Four seconds later he has a paragraph. It is grammatical, it is confident, and it quotes a federal standard.
That paragraph is the problem, and not for the reason most people expect.
The failure mode is fluency, not nonsense
Hallucination takes the attention because it is visible. An answer that invents a standard number, or describes a permit class that does not exist, gets caught by the first person who reads it. Obvious wrongness is self-limiting.
The answer that does damage is correct as a statement of law and wrong as an instruction on that site. In the United States, the federal lockout standard does not tell anyone how to isolate a specific pump. 29 CFR 1910.147(c)(4)(i) requires that procedures “shall be developed, documented and utilized for the control of potentially hazardous energy.” The regulation delegates the actual answer to a document the employer wrote. An assistant that retrieves the regulation and paraphrases it has retrieved the instruction to have a procedure. It has not retrieved the procedure.
This is the structural point, and it holds well beyond lockout. A regulation is a floor. American law says so explicitly: under section 18(c)(2) of the OSH Act, a state plan must adopt and enforce standards that are “at least as effective” as the federal ones. Twenty two state plans cover private sector workers, and seven more cover state and local government workers only. The federal text is the minimum around which everything else is built upward.
Most site procedures sit above that minimum too, for reasons that are rarely written down next to the rule: an incident in 2011, an insurer’s condition, a corporate standard harmonised across three countries, a plant manager who decided the permit threshold was too loose. None of that is in the Code of Federal Regulations. All of it is in the site procedure. An assistant retrieving public regulatory text will pull the worker down to the floor with correct grammar and no signal that it has done so.
Jurisdiction is the second trap, and it is quieter
Ask an assistant what the rules require for working in heat. In California, 8 CCR 3395 is in force for outdoor work and requires a written prevention plan, drinking water, access to shade, and cool-down rest, with additional high heat procedures from 95 degrees Fahrenheit. At federal level, the heat rule was published as a proposed standard on 30 August 2024 and is still in rulemaking. Two answers, both accurate, separated by a state line that the person holding the tablet cannot see in the retrieval index.
Now move the same question across a national border. The unit system changes, the exposure limits change, the permit regime changes, and the document the assistant surfaces may be a translation of a standard that was never adopted where the worker is standing. The failure mode does not get better outside the United States. It gets harder to detect, because the retrieved text still looks authoritative.
There is some published evidence that retrieval does not solve this class of problem. A preregistered evaluation of retrieval based legal research tools, published in the Journal of Empirical Legal Studies in 2025, found the tools produced hallucinated or unsupported output between 17 and 33 percent of the time despite grounding claims. Treat that carefully. It is US case law research, not safety procedure lookup, and the error type measured there is not the error type described here. It is not a benchmark for your site. It is a reason to measure your own.
Where these tools do earn their place
None of this is an argument against putting retrieval in front of workers. It is an argument about what you point it at.
Index your own controlled documents, not the open internet and not a regulatory corpus. Scope retrieval to the current revision of your procedures, permits, safety data sheets, isolation lists and drawings, with revision control enforced at the index so that a superseded document cannot be returned.
Return the document, not the paraphrase. The most useful output is a pointer with the passage shown: procedure reference, revision number, section, and the clause visible so the reader judges it. A worker who is handed the right page in four seconds has gained something real. A worker who is handed a summary has lost the audit trail.
Use it for the meta questions, which are most of the questions. Which permit applies. Who is the authorised person. Where is the current form. When did this last change. These have document-shaped answers and no regulatory ambiguity.
Use it to draft for a reviewer. Job safety analyses, toolbox talk outlines, a first pass at summarising a data sheet, all going to a named competent person who signs. Drafting is not deciding.
The diagnostic
Pull 20 questions from your own procedures where the site requirement sits above the regulation, or where the answer changes by state, country or unit. Put them to the assistant exactly as a worker would type them on shift, with no follow-up prompt and no expert rephrasing. Then count: how many answers came back to the regulation rather than to your written procedure? Zero or one means retrieval is genuinely scoped to your controlled documents and you have a tool worth deploying. Two to five means the index is leaking public regulatory text into operational answers and needs to be re-scoped before it goes in front of anyone. More than five means the system is not answering for your site at all, and the demo that impressed the steering group was measuring fluency.
The duty does not transfer
There is a legal point underneath the operational one. In the United States, 29 CFR 1910.1200(h)(1) states that “employers shall provide employees with effective information and training on hazardous chemicals in their work area at the time of their initial assignment.” Paragraph (h)(3)(iv) requires that the training cover “the details of the hazard communication program developed by the employer.” The duty sits on the employer and the content is the employer’s own programme. An answer engine is a reference source. It cannot discharge a training obligation, and no amount of usage data will make it look like one to a compliance officer. In state plan states the same duty applies through a standard that must be at least as effective.
For the governance framing, the NIST AI Risk Management Framework (AI RMF 1.0) is voluntary and organised around four functions: govern, map, measure and manage. The companion Generative AI Profile, NIST AI 600-1, names confabulation as a distinct risk category. The word doing the work in that framework is measure, and measure means measure in the deployed context. The deployed context is a worker on shift with a question, not a curated demo with a prompt engineer in the room.
Compliance is not safety, and a system that answers to the compliance floor is not answering the question that was asked. The 20 questions cost an afternoon to write. Write them before the business case, attach the count to the procurement file, and require the vendor’s score on your questions rather than theirs.