Black Swan events in process safety
≈ 11 min read · 2,173 words
In a well-known experiment people watch a video of a basketball being passed around, and their only task is to count the passes made by the players in white. Meanwhile someone in a gorilla suit walks into the frame, and most of the counters fail to notice. They are not inattentive: they were looking for exactly what they had been told to look for. The same blind spot operates in a plant if safety means nothing more than eliminating the hazards listed in advance.
A Black Swan is an extremely high-impact event that past experience cannot predict, yet which looks explainable in hindsight.
It does not follow from our existing models, so frequency-based forecasting cannot be applied to it. The phenomenon is common in systems that live in “Extremistan”, where a single outlier can dominate the whole, as opposed to “Mediocristan”, where the extreme is negligible. The defence is not prediction but preparedness: robustness and redundancy.
Figure 1 — in Mediocristan the bell curve is bounded and the mean is meaningful; in Extremistan the outlier Black Swan of the fat-tailed distribution dominates.
Who is this for?
Section titled “Who is this for?”For those who work with rare but severe events and with the resilience of the system: plant manager · process engineer · HSE specialist · process safety engineer · reliability engineer · shift supervisor.
Learning objectives
Section titled “Learning objectives”After reading this article you will be able to:
- distinguish a low-probability event from an unimaginable one;
- decide whether a phenomenon lives in Mediocristan or in Extremistan, and align your metrics accordingly;
- show where the limit of frequency-based risk assessment lies in a process plant;
- list the tools of preparedness, and recognize when an event is NOT a Black Swan.
In brief
Section titled “In brief”- Black Swan = unimaginable, not merely improbable. It does not follow from our models; it can only be explained after the fact.
- In Mediocristan the mean and the trend are meaningful; in Extremistan a single outlier dominates, the mean misleads, and Black Swans occur at every turn.
- The “smart person in Extremistan” does not predict but adapts in advance, and that requires convertible knowledge.
- Process safety consequence: frequency-based risk assessment (HAZOP/LOPA) can underestimate rare, severe events, which is why resilience is needed (robustness and redundancy).
The image of the “black swan” is itself misleading, since we can perfectly well imagine a black swan. The task is to prepare not for the improbable but for the unimaginable.
What is a Black Swan event, and where does it come from?
Section titled “What is a Black Swan event, and where does it come from?”The concept comes from the work of Nassim Nicholas Taleb. By his definition, Black Swans are extremely rare and hard-to-predict events whose consequence is disproportionately large, and for which “human nature manufactures an explanation after the fact”.
Taleb’s own life is instructive too: after the collapse of his home country, Lebanon, he learned the toolkit of finance at Wharton, and then applied it to its opposite in his own firm. For years he produced a daily loss, but on 11 September 2001 he gained a great deal, and in the 2008 crisis he became a multi-billionaire.
The Hungarian treatment comes from László Mérő’s ELTE lecture (International Lean Summit, 2013). Mérő sharpens the concept: these are “not low-probability events” but “unimaginable” ones. Three characteristics hold together: extreme impact, retrospective explainability (obvious afterwards, though it was not so in advance) and prior unpredictability.
Typical examples: 11 September 2001, the 2008 financial crisis, the collapse of the Soviet Union. The recurring error is the claim that “a crisis of this size does not happen even once in ten thousand years” — while 1987, 1998, 2001 and 2008 all did happen. The phenomenon has intensified because accumulation, a changing environment, growing complexity and the Matthew effect act together.
Mediocristan and Extremistan: two worlds, two logics
Section titled “Mediocristan and Extremistan: two worlds, two logics”In Mediocristan the phenomena are bounded, so the mean describes the typical case; in Extremistan a single outlier can dominate, so the mean is misleading. The difference shows up even in a quiz:
| Question: what is the average… | Answer | World |
|---|---|---|
| of Hungarians taller than 2 metres? | 204 cm | Mediocristan |
| of people over 90 years old? | 93 years | Mediocristan |
| of Hungarians with wealth above HUF 5 billion? | HUF 22 billion | Extremistan |
| of companies with capital above USD 5 billion? (in August 2012 there were 761 of them) | USD 27 billion | Extremistan |
The two worlds differ not only in the mean, but also in what behaviour pays off within them:
| Characteristic | Mediocristan | Extremistan |
|---|---|---|
| Scientific knowledge | plenty of it, the workings are predictable | little, and what there is (chaos theory, fractal theory, scale invariance) is precisely about unpredictability |
| Useful trait | rule-following | scepticism |
| Black Swan | truly only once in ten thousand years | at every turn |
| History | flows slowly | advances in sudden jumps |
| Guiding principle | “water is the master” | “imagination is the master” |
As a good approximation, the dividing line is that what man has made easily belongs to Extremistan, while what nature makes mostly belongs to Mediocristan. An engineered system is therefore man-made, that is, prone to Extremistan behaviour. Everyday life apparently runs in Mediocristan, but it can switch over at any moment, which is why dual thinking is needed. The systemic error is applying a Mediocristan tool to an Extremistan problem.
The “smart person in Extremistan”: convertible knowledge
Section titled “The “smart person in Extremistan”: convertible knowledge”There are two kinds of smartness, and the two are not interchangeable:
- The person who is smart in Mediocristan possesses specialized knowledge and predicts the future well; techniques can be built from their smartness. For centuries this was the engine of progress.
- The person who is smart in Extremistan sees clearly that the future cannot be predicted even approximately, and yet is able to adapt in advance. This calls for knowledge convertible to anything: thinking capable of shifting perspective, instead of rigid expertise.
Taleb’s advice, the closing thought of the lecture: “Let us choose wisely when to be stupid” — that is, when to let go of our ingrained models. Resilience engineering calls this adaptive capacity.
How do you defend against it? (preparedness, not prediction)
Section titled “How do you defend against it? (preparedness, not prediction)”What protects against a Black Swan is not more accurate forecasting (by definition it cannot be forecast), but preparedness:
- Acknowledge the limit: the timing and the form of the extreme event are unpredictable.
- Build robustness and redundancy. This protects not against a scenario but against the unknown (robustness and redundancy).
- Avoid over-optimization, which eliminates the buffer: “there is no optimization, but there is common sense”.
- Keep it simple: simplicity is a value in itself, because it hides fewer latent failure modes.
- Institutionalize scepticism: ask regularly, “what if the unimaginable happens?”
- Stay open to positive Black Swans as well: the extreme waves that come along can be ridden, and this is scale-invariant. Still, the most can be achieved by generating a wave of your own.
What do Mediocristan and Extremistan have to do with process safety?
Section titled “What do Mediocristan and Extremistan have to do with process safety?”In the process industry the Black Swan is the typical “low-probability, high-consequence” event: a never-modelled combination of failures, a rare external event, a domino effect. Frequency-based risk assessment (HAZOP, LOPA and SIL) can underestimate these or fail to catch them at all, because it is built on past data.
The established approach is to define all possible risks in advance, and to design procedures which, followed exactly, avoid every known hazard. This safety by protocol is excellent in principle; its single critical flaw is the illusion of knowledge: it assumes that every risk can be known in advance. The gorilla experiment shows exactly this. A HAZOP or FMEA team is not given such a narrowing, and is still vulnerable to rare, unexpected events, such as the magnitude 9.0 Japanese earthquake and tsunami of 11 March 2011: citing Taleb, it is precisely the rare and unexpected events that are the truly dangerous ones. The textbook process-industry example is Deepwater Horizon, which the human reliability literature discusses as an example of Taleb’s Black Swan concept.
This is why process safety thinking is complemented by resilience: independent, redundant protection layers and sceptical review (“what-if” alongside HAZOP). Daily OEE and the small stoppages move in Mediocristan, whereas the major safety events move in Extremistan, where a single event can dominate the whole year’s loss. KPI-based daily management on its own does not protect against rare, severe events.
Common mistakes
Section titled “Common mistakes”Most of the errors come from carrying Mediocristan reflexes over into an Extremistan situation:
- “Improbable” ≠ “unimaginable”. The Black Swan is not rare; it is what the model leaves out.
- The mean as a safety yardstick. The average of accident statistics can hide the catastrophic tail.
- Over-optimization in the name of “efficiency”. The reserve you eliminated is missing precisely when it would be needed.
- Clinging to a single model. The absence of a shift in perspective blinds you to situations that do not fit the procedures.
- Confirmation bias, in Taleb’s terms “the error of confirmation”: we look only for information that confirms our expectations. This turns the review into self-justification.
When NOT to use it?
Section titled “When NOT to use it?”- For known, modellable risk. It does not replace the HAZOP / LOPA and SIL system, and it is no excuse for neglecting known hazards.
- For recurring events describable with frequency data (seal leaks, small stoppages). These call for a Mediocristan tool: statistics, trend, OEE.
- As an after-the-fact excuse. If the event appeared in the risk assessment, or there was a weak signal for it, it is not a Black Swan but a missed protection layer.
- Nor is the reverse an absolute. No amount of preparedness prevents every problem; the goal is resilience.
Take it home (keys)
Section titled “Take it home (keys)”- Ask which world you are in. If a single event can dominate the annual result, the mean is not a metric.
- Sharpen resilience, not prediction. Redundancy protects not against a scenario but against the unknown.
- The buffer you saved is a risk. Before you withdraw a reserve, say out loud which event the protection ceases against.
- Simplicity is a safety argument: fewer latent failure modes.
- Build scepticism into the routine, and watch the weak signals. You cannot see the Black Swan coming, but you can see a protection layer weakening.
Self-test
Section titled “Self-test”- To which event can frequency-based forecasting be applied: to the low-probability one or to the Black Swan? Why?
- Which trait pays off in Mediocristan, and which in Extremistan? Why exactly that one?
- Give an example from your own plant of a Mediocristan and an Extremistan phenomenon, and say which tool suits which.
Answer key: 1) To the low-probability one: it is inside the model. Not to the Black Swan; only preparedness works on that. 2) In Mediocristan rule-following, in Extremistan scepticism, because a rule fixed in advance is blind precisely to the unmodelled case. 3) For example small stoppages (statistics, trend), and catastrophic release (redundancy, barrier review).
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”Although the Black Swan cannot be predicted, the weak signals and the state of the protection layers can be, and recording exactly these is the job of the shift log (OPEREX): near-misses (near-miss), lost reserves and barrier weaknesses, documented shift by shift, do not go unnoticed. Recording near-misses and deviations gives an auditable trail (ISO 45001 §10.2: management of incidents and nonconformities).
Terminology (HU / EN)
Section titled “Terminology (HU / EN)”| Hungarian | English | Note |
|---|---|---|
| Fekete Hattyú | Black Swan | an unimaginable, high-impact event |
| Átlagisztán | Mediocristan | bounded phenomena |
| Extremisztán | Extremistan | the outliers dominate |
| Vastag farkú eloszlás | Fat-tailed distribution | the rare extreme is not negligible |
| Skálafüggetlenség | Scale invariance | the pattern is the same at small and large scale |
| Megerősítési torzítás | Error of confirmation | only information confirming the expectation |
| Konvertálható tudás | Convertible knowledge | flexible, transferable capability |
| Adaptív kapacitás | Adaptive capacity | handling situations beyond the rules |
What is the difference between a low-probability event and a Black Swan event?
The low-probability event is inside our models, it just occurs rarely; the Black Swan, by contrast, does not follow from our models, so on the basis of earlier knowledge it is not “improbable” but “unimaginable”.
Why is the mean misleading in Extremistan?
Because a single outlier (one super-rich person or one catastrophic outage) distorts the average so much that it no longer represents the typical case, and no forecast can be based on it.
Why do statistical forecasts not protect against Black Swans?
Because forecasting is based on extrapolating past patterns, whereas the Black Swan by definition does not follow from the past. The defence is therefore preparedness: robustness, redundancy and sceptical review.
What does "convertible knowledge" mean?
A flexible, transferable problem-solving capability that can be applied to almost any new situation, instead of narrow expertise.
Related concepts
Section titled “Related concepts”robustness and redundancy · near-miss · HAZOP · LOPA and SIL · muda
Next step
Section titled “Next step”- robustness and redundancy — the engineering toolkit of preparedness.
- near-miss — collecting the weak signals.
References / further reading
Section titled “References / further reading”- Nassim Nicholas Taleb: The Black Swan: The Impact of the Highly Improbable. Penguin Books, 2007.
- László Mérő: Preparing for the unimaginable. ELTE, International Lean Summit, 2013.
- Carl S. Carlson: Effective FMEAs. Wiley, 2012, chapter 2.4: the critique of safety by protocol.
- Designing for Human Reliability: Human Factors Engineering in the Oil, Gas and Process Industries, 2015.
In practice
Operational excellence for industrial operations — adopted module by module, starting with the digital shift log
Learn more: Incident investigation →