The AI That Writes Your Incident Report Learned Root Cause From the Reports Before It
AI tools that draft post-incident investigation reports are pattern-matched against a corpus of past reports that already over-attributes causation to operator error, so faster drafting will scale that bias, not correct it.
A storage tank at a refinery kept overflowing. Not once. Over several years, more than once, spilling hydrocarbons onto the ground each time. Management’s read was simple: the operators weren’t paying attention. They disciplined the workers involved, retrained them, tightened the rules, and warned that the next overfill would not be tolerated. It happened again anyway. When someone finally looked past the operators, the answer was in the hardware: the tanks had no high level alarms to warn the control room before they overtopped. The U.S. Chemical Safety and Hazard Investigation Board (CSB) has told this story to make one point: blaming the person in front of the failure is often the investigation ending too early, not the investigation working.
That CSB account is now two and a half decades old, and the pattern it describes has not gone away. It is worth revisiting now because a new tool is entering the incident investigation workflow, and it is trained on exactly the corpus that produced stories like the tank farm.
What the tool actually does
AI drafting assistants aimed at EHS teams promise to compress the slowest part of incident investigation: turning witness interviews, timeline notes, and photos into a structured report with a suggested root cause and draft corrective actions. That is a real time cost worth cutting. A five-whys writeup that used to eat an afternoon can now produce a first draft in minutes.
The part worth slowing down on is where “suggested root cause” comes from. These tools do not investigate the plant. They pattern-match language: witness statements and timeline facts against the phrasing and causal structure of thousands of historical reports the model was trained or fine-tuned on. If most of the reports in that training set describe a similar sequence of events and land on “operator failed to follow procedure” or “inadequate attention to task,” the model has learned that this is the shape of a normal answer. Give it a similar sequence of events and it will tend to reach for the same shape.
The bias is documented, not hypothetical
This is not a new fear invented for AI. It is a known, named failure mode of human investigation that predates any software. Investigators researching cause attribution in workplace incidents have described hindsight bias and outcome bias: once an outcome is known, it gets read backward as having been more foreseeable and the people involved as having been more careless, than the facts available at the time actually supported. A peer-reviewed review of this literature in Applied Ergonomics lays out how these biases push investigators toward individual decisions and missed opportunities as the explanation, and away from the equipment, procedure, and workload conditions that were also in the room at the time (Hutchinson et al., 2022, via PubMed).
The CSB’s own program staff have made the same point about their field directly. Writing about the agency’s early chemical incident investigations, a CSB program analysis officer noted that attributing an accident to human error is “about as helpful as listing gravity as the cause of a fall,” quoting process safety engineer Trevor Kletz, and argued that the pull toward blaming the individual closest to the failure is strong precisely because it feels more resolved than tracing multiple management system failures back through the organization (CSB, 1999). The Center for Chemical Process Safety’s Guidelines for Investigating Process Safety Incidents exists in large part because “operator error” as a stopping point, rather than a starting point for asking why the error was possible, was common enough across the industry to warrant a standing methodology correcting for it.
Put those two things together. If the historical corpus of investigation reports carries a documented lean toward individual attribution, and an AI drafting tool is trained on that corpus, the tool has no independent way to know the lean is a bias rather than a pattern worth repeating. It was never shown the counterfactual reports, the ones where a company dug past the operator and found the missing alarm, the understaffed shift, or the procedure nobody could actually follow. It only knows what got written down.
Monday morning check
Pull your last five closed incident reports and count how many corrective actions are aimed at a person (retrain, discipline, coach, remind) versus how many are aimed at a system (redesign, add an interlock, change staffing, fix a procedure). If an AI tool drafted or suggested any of those root causes, ask it, or the vendor, what training data or reference reports shaped that suggestion. **If nobody can answer that question, the tool's root cause is a guess dressed up as analysis.**
Faster is not the same as better
None of this means the drafting tools are useless. A well-supervised tool that assembles a timeline from scattered witness notes, flags inconsistencies between statements, or drafts the boilerplate sections of a report can free an investigator’s time for the harder work: asking why the procedure was unworkable, why the alarm didn’t exist, why the shift was short-staffed. That is a legitimate, bounded use.
The failure mode is treating the tool’s suggested root cause as the analysis rather than a first draft to be argued with. A report that reaches a conclusion faster is not automatically a report that reaches a better conclusion. Speed only helps if the thing being sped up was sound to begin with. If the training corpus systematically stops the causal chain at the worker, an AI tool will draft that same stopping point with more confidence and less friction than a tired investigator typing at midnight, which is arguably worse, because confidence reads as rigor.
What would actually tell you the tool is working
This is genuinely early. There is no published, peer-reviewed study yet directly measuring whether AI-assisted investigation reports show a different attribution ratio than human-only reports covering comparable incidents; that comparison has not been done publicly, and anyone who tells you it has should be asked for the citation. Until it is, the honest position is to hedge: treat AI-suggested root causes as a hypothesis generated from historical pattern, not a finding, and check it against the same standard CCPS guidance recommends for human investigators, that the labeled root cause be a system or management factor whose correction would plausibly have prevented the event, not a restatement of what the worker did wrong. If your AI-assisted reports keep landing on human error at roughly the same rate your pre-AI reports did, the tool has not changed your investigation practice. It has only changed how long it took to write the same conclusion down.
The tank farm story ends the way it does because someone finally asked why attentive, experienced operators kept making the same “mistake.” That question, not the speed of the writeup, is what root cause analysis is for. A tool that answers faster without being asked to justify its answer is not saving that step. It is skipping it.