Robustness and redundancy
≈ 11 min read · 2,101 words
The spare wheel lies at the bottom of the boot for years, producing nothing, until one single night it pulls you out of trouble. In a plant the same thing is the spare pump, the second pressure transmitter or the final safety valve. The question is not whether it is worth it, but which spare actually protects.
Robustness is staying operable under unforeseeable, extreme disturbance. Its source is redundancy: a spare that fails the same way protects nothing.
Figure 1 — the over-optimized serial chain collapses on a single failure; of the redundant layers one can be lost and the system still works.
Who is this for?
Section titled “Who is this for?”For those who decide on the spares of critical equipment, or who live with the state of the protection layers every day: plant manager · process engineer · shift supervisor · HSE / process safety specialist · reliability engineer · instrument technician and maintenance staff.
Learning objectives
Section titled “Learning objectives”After reading this article you will be able to:
- distinguish identical from diverse redundancy, and name the common cause failure (CCF);
- decide with three questions whether a spare is real protection or a covered-up problem;
- recognize where redundancy is not the right answer.
In brief
Section titled “In brief”- Robustness = survival against the unknown, not accurate prediction; its source is redundancy.
- Over-optimization makes you fragile: the buffer you removed is exactly what is missing in the crisis.
- A spare is only a spare if it is diverse: with identical design and maintenance a common cause failure (CCF) is looming.
- The independent layers of LOPA are the institutionalized form of redundancy: the critical spare is not muda.
Why it matters (the stakes)
Section titled “Why it matters (the stakes)”The weakening of robustness makes no noise. After a spare is removed or an interlock is bypassed, the system behaves the same way for months, and the decision looks like clever cost cutting. The bill arrives when the rare event occurs, and exactly the layer is missing that nobody missed.
What is it, and where does it come from?
Section titled “What is it, and where does it come from?”The idea comes from N. N. Taleb’s Black Swan theory (The Black Swan, 2007); its Hungarian exposition was given in László Mérő’s 2013 Lean conference talk (“Preparing for the Unimaginable”, Lean Summit 2013). For the conceptual treatment of the predictable (Mediocristan) and the unpredictable (Extremistan) domain, see the black swan event article. In Extremistan there is no prediction and no optimization, “there is common sense, though”, and what is needed is not yet another methodology: even the word “methodology” is painfully Mediocristanean. What is needed is redundancy and simplicity.
The two domains reward different behaviour:
| In Mediocristan (normal operation) | In Extremistan (the rare, extreme domain) |
|---|---|
| rule following and compliance is the useful trait | skepticism is the useful trait |
| optimization, forecasting, sophisticated techniques | redundancy, simplicity, adaptability |
In daily operation following the procedure is the right answer, not improvisation. Skepticism is added for the extreme domain, not instead of it: what happens if the unimaginable does occur after all?
What kind of redundancy are we talking about?
Section titled “What kind of redundancy are we talking about?”The functional safety standard (IEC 61511 / MSZ EN 61511) gives precise words for what the word “spare” blurs:
- Redundancy: extra devices performing the same function beyond what is necessary, to increase reliability and availability; it can be implemented with identical elements (identical redundancy) or with different ones (diverse redundancy).
- Diversity: the same function realized in several ways, on principles that differ from one another.
- Fault tolerance: the unit performs the required function even in the presence of faults. This is the engineering definition of robustness.
- Common cause failure (CCF): a failure that causes identical faults in two or more separate channels of a multi-channel system.
CCF is the standard name of apparent redundancy: a single cause (a shared power supply, instrument air, logic solver or maintenance practice) takes the “independent” layers out together.
Figure 2 — with identical design and maintenance the spares fail together (CCF); real fault tolerance comes from the diverse spare.
Why do redundancy and robustness matter in unpredictable (Extremistan) systems?
Section titled “Why do redundancy and robustness matter in unpredictable (Extremistan) systems?”Because the rare, severe event cannot be predicted, only survived, and survival does not come from the accurate model but from having several independent chances to stop it. Process safety institutionalizes this: LOPA prescribes several independent protection layers (BPCS control, critical alarm and operator intervention, SIS/SIF, physical protection, consequence mitigation), because the failure of a single layer must not lead to a catastrophe. The more truly independent IPLs stand in the chain, the smaller the residual risk.
When is it a spare, and when is it a covered-up problem?
Section titled “When is it a spare, and when is it a covered-up problem?”Not every spare protects, and the difference can be settled with three questions. In the process industry it is not inventory that accumulates, but technology, engineers, extra equipment and operational support; according to the foundational work of process-industry Lean (Floyd: Liquid Lean), an unnatural accumulation of resources is usually the sign of an unsolved problem.
In a chemical plant (anonymized case) at some services there stood not one but two spare pumps beside the main machine. One spare is a normal contingency; two are already suspicious. Because of the high start-up viscosity the technicians started the spare motors every week so the rotor would not get a flat spot: with this they exposed them every week to the most severe stress they would ever experience, all of them with the same practice. That is why, shortly after the main pump failed, both spares failed too. The solution was not a new machine but turning the shaft by hand.
This is where the test comes from. A spare is real redundancy if:
- its failure mode differs from that of the main element;
- its maintenance and testing practice differs;
- the HAZOP/LOPA has credited it as an IPL, so it is independent and auditable.
If the answer to any of these is no, the item is not redundancy but a covered-up problem. This resolves the apparent conflict between Lean and process safety: where there is no protective function, the cause of the accumulation must be eliminated (waste), but at a critical protection point the spare must be kept and made diverse, because there the measure is not cost efficiency but the credited IPL and availability.
Capability reserve works the same way: a process with greater process capability (excess capability) is more robust, because it tolerates some special cause variation without producing off-spec product.
Process industry context
Section titled “Process industry context”In Seveso plants (crude oil processing, chemical plant) robustness takes shape in the spares of critical equipment (a standby safety pump, redundant instrumentation, 2oo3 voting logic in the SIS) and in simple, transparent procedures.
Redundancy, however, is not only hardware. According to the human factors literature, people are the most resilient and robust element of any sociotechnical system, capable of finding a solution even to an unforeseen situation; conversely, the expectation placed on human reliability may be lower if the other elements of the system can compensate. A shift with sufficient headcount, trained and not overloaded, is therefore itself a protective reserve: the only one that can improvise.
Hands-on (doable in an hour)
Section titled “Hands-on (doable in an hour)”List the three most critical protection layers (IPLs) of your plant, and write down what each shares with its spare: power supply? instrument air? logic solver? maintenance practice? Wherever there is even one shared element, the twofold redundancy on paper is a single one in reality. Take the result into the next HAZOP or LOPA review.
Common mistakes
Section titled “Common mistakes”Most mistakes come not from the absence of redundancy but from misunderstanding it.
- Cutting a critical spare as “waste” → first put the three questions to it.
- Identical maintenance on every spare → the faulty practice takes them all out at once.
- Complexity in the name of “safety” → complicated protection brings new, hidden failure modes.
- A spare out of service stays hidden → take it into the shift handover, otherwise fragility grows unnoticed.
When NOT to use it? (the limits of the method)
Section titled “When NOT to use it? (the limits of the method)”Redundancy protects against the residual, non-eliminable uncertainty: for a known and fixable fault the fix is the cheaper and more robust answer, and the spare then only adds cost and complication.
| Situation | Why redundancy is not the answer | The right answer |
|---|---|---|
| Design error (undersizing, wrong material) | you duplicate the faulty design | redesign, MOC |
| A known, unfixed root cause | the spare covers it, the problem remains | root cause analysis, near-miss investigation |
| Missed maintenance or test | the unmaintained spare will not start | test interval, proof test |
| Common cause failure (CCF) | an identical layer gives no new chance | diversity, separation |
Take it home (keys)
Section titled “Take it home (keys)”- Count the independent layers, not the pieces (two identical spares are one chance).
- Ask the spare: different failure mode? different maintenance? credited IPL? Three yeses are needed.
- Where there is no protective function, accumulation is a symptom: look for the unsolved problem behind it.
- Simplicity is cheap robustness: every new layer also brings a new, hidden failure mode.
- Keep a record of the disabled layer: what has no entry, nobody knows is missing.
Self-test
Section titled “Self-test”- What is the difference between identical and diverse redundancy, and which one protects against CCF?
- In the spare pump case, what made the redundancy apparent only?
- Name two situations in which a further protection layer does not increase robustness.
How does this show up in digital practice?
Section titled “How does this show up in digital practice?”The current level of robustness can be read from how many protection layers are disabled right now. The principle is the same in any well-run plant: a time-limited register of the interlock overrides (MOS/POS), the due dates of the trip tests, a status view of the spares out of service and the management of change workflow together show how many real layers the system stands on today.
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”Recording a spare that is out of service (for example a temporarily inoperable safety pump or SIF) and the interlock override (MOS/POS) at the shift handover is the essence of the practice under IEC 61511: the disabled IPL is registered, time-limited and verified back. In the shift log (OPEREX) this can be done auditably; if the loss is handled as a nonconformity, the entry also leaves an audit trail under ISO 45001 §10.2.
Terminology (HU / EN / JP)
Section titled “Terminology (HU / EN / JP)”| Hungarian | English | Japanese | Note |
|---|---|---|---|
| Robusztusság | Robustness | ロバスト性 | operability under extreme disturbance |
| Redundancia | Redundancy | 冗長性 | a device of the same function beyond what is necessary |
| Azonos / diverz redundancia | Identical / diverse redundancy | 同一冗長・多様冗長 | with identical, respectively different elements |
| Diverzitás | Diversity | 多様性 | the same function on a different principle |
| Hibatűrés | Fault tolerance | 耐故障性 | it works even with a fault present |
| Közös okú hiba | Common cause failure (CCF) | 共通原因故障 | two “independent” layers fail from one cause |
| Független védelmi réteg | Independent Protection Layer (IPL) | 独立防護層 | the basic element of LOPA |
| Retesz-feloldás | Maintenance / Process Override Switch (MOS/POS) | インターロック解除 | the register of the disabled IPL |
Why does over-optimization make a system fragile?
Because in the name of efficiency it eliminates exactly the seemingly superfluous spares, so no buffer is left when an unforeseen extreme event occurs.
What is the difference between identical and diverse redundancy?
Identical redundancy is built from identical elements, diverse redundancy from different elements or on a different principle. Only the diverse one protects against common cause failure (CCF), when a single cause (a shared power supply, a shared maintenance practice) takes out all the channels at once.
Does redundancy contradict Lean waste reduction?
Only apparently. The critical safety spare is not waste, because it gives robustness against rare, severe events; a spare that protects nothing, however, is mostly the cover of an unsolved problem, and there the Lean logic is the right one.
Why is simplicity a value in robustness?
Because complexity brings mutual dependences, and with them new, hidden failure modes. A simple system is less fragile, and it reduces the chance of extreme events arising from human error.
Related concepts
Section titled “Related concepts”black swan event · LOPA and SIL · HAZOP · near-miss · management of change · muda
Next step
Section titled “Next step”- black swan event — the conceptual exposition of Mediocristan and Extremistan.
- LOPA and SIL — how many independent protection layers are enough, and when a SIF is needed.
- management of change — how not to lose your redundancy on a “minor” modification.
References / further reading
Section titled “References / further reading”- IEC 61511 / MSZ EN 61511-1,-3: Functional safety in the process industry. The canonical definitions of redundancy, diversity, fault tolerance, CCF and the interlock override (MOS/POS).
- Nassim Nicholas Taleb: The Black Swan. Random House, 2007 — the foundational work on high-impact, unpredictable events.
- Raymond C. Floyd: Liquid Lean. Productivity Press, 2010 — the Lean reading of resource accumulation in the process industry.
- Karl E. Weick and Kathleen M. Sutcliffe: Managing the Unexpected. Jossey-Bass — high reliability organizations (HRO).
In practice
The daily state of the protection layers, the spares that are out of service and the interlock overrides (MOS/POS) can be recorded auditably in the shift log (OPEREX), so the actual availability of the credited IPLs is traceable shift by shift.
Learn more: Incident investigation →