Root cause analysis (RCA)
≈ 15 min read · 3,034 words
Root cause analysis (RCA) is a structured, team-based method that traces the causal chain back from a failure to the root cause, and then eliminates it.
Its founding premise is that if you treat only the surface symptoms, the problem will almost certainly come back; the durable fix is identifying and eliminating the root cause. The output is not an explanation but a verified corrective action, which turns a reactive “firefighting” culture into a forward-looking, reliability-building one. Applied to equipment failures it is also called root cause failure analysis (RCFA).
Figure 1 — the symptom and the root cause are not the same. If you treat only what is visible on the surface, the problem returns; RCA digs down to the root cause, where it can be eliminated for good.
Who is this for?
Section titled “Who is this for?”This article is for those who work with recurring failures and their investigation: plant manager · process engineer · maintenance and reliability engineer · shift supervisor · HSE specialist · quality · CI coordinator.
Learning objectives
Section titled “Learning objectives”After completing this module you will be able to:
- distinguish between the symptom, the problem and the root cause
- work through the six steps of root cause analysis
- decide which event needs a full investigation and which only a quick report
- put the tools in order (fishbone, fact tree, five whys), and place RCA (reactive) relative to FMEA (proactive)
In brief
Section titled “In brief”- RCA looks for the root cause, not the symptom; otherwise the failure (and the needless production loss) recurs.
- Problem = a deviation from the desired state or from the norm; symptom = the state caused by the problem; the series of events leading to it is the problem chain; root cause = the fault from which the whole chain of effects originates.
- Systematic investigation: timeline, sequence of events, documented evidence, teamwork; not brainstorming for its own sake.
- Six steps: define the problem → collect data and evidence → identify the causal factors → develop solutions → implement → track the effectiveness.
- Not every failure is investigated to full depth: the level is decided by severity.
- The tools have an order: fishbone (collecting candidates) → fact tree (logical linking) → five whys (closing and deepening).
- A reactive method (after the failure), as opposed to the proactive logic of FMEA; both serve the same end: prevention.
What it is and why it matters
Section titled “What it is and why it matters”Unreliable asset performance is a threat in most operating environments, and a reactive culture “fixes” the same failure again and again. RCA eliminates the problem at its source, so over time the frequency of problems falls. It often also uncovers organizational or cultural roots (incentives that set production against maintenance, insufficient empowerment, rigid boundaries between trades), the removal of which reaches beyond local optimization; this is why RCA is regarded as equivalent to the kaizen improvement process.
What is root cause analysis (RCA), and what are its six steps?
Section titled “What is root cause analysis (RCA), and what are its six steps?”RCA works backwards from the failure to the triggering root cause, and then eliminates it. It is not a single, rigid methodology, but the general process that can be regarded as universal consists of six steps.
Figure 2 — the six steps; the process is iterative: repeat it until the problem has genuinely gone away.
- Define the problem (the failure). What exactly is the deviation from the desired state or from the norm? Without a good problem statement the root cause cannot be identified either.
- Collect data and evidence about the difficulties that contributed to the problem: what happened, when, and in what order. This is where the timeline / sequence of events and the problem chain are recorded.
- Identify the possible causal factors. What contributed, and how are they connected? This is the heart of the analysis: the RCA tools take you from here to the actual root cause(s), while candidates unsupported by evidence are ruled out.
- Develop solutions and recommendations. The countermeasure that prevents recurrence; among equivalent alternatives, the simplest and cheapest.
- Implement the recommendations, with an owner and a deadline.
- Secure the effectiveness: track the solutions that were implemented. The established schemes review the actions again in a separate review 12–18 months later, because a single intervention rarely eliminates the event: the process is iterative. Carry the lesson over to sister equipment and to the other sites as well.
Which failures should be investigated at all?
Section titled “Which failures should be investigated at all?”Not every event needs a full RCA: the depth of the investigation is set by severity. Established practice uses a harmonized consequence table on a five-point scale (1 minor: 1–10 kUSD or 0–6 hours of loss; 3 serious: 0.1–1 mUSD or 12 hours–3 days; 5 extensive: above 10 mUSD and 15 days), and the highest individual consequence determines the ranking. The values are illustrative orders of magnitude; the bands are set by the organization.
| Severity | Who leads the investigation | Mandatory output |
|---|---|---|
| Low | the leader of the affected asset team | quick report |
| Medium | an investigator appointed at site level, with a cross-functional team | quick report + short summary (“5-pager”) |
| High | an appointed investigator, with an external or subject-matter expert involved | quick report + summary, with a group-level review |
Inputs: production loss, HSE incidents, process concerns, system defects; from these a continuously maintained top list is built on the basis of total business and HSE impact. The guiding principle: better to do excellent work on the most important things than mediocre work on many.
The RCA toolbox
Section titled “The RCA toolbox”The root cause is typically the resultant of several contributing factors, which is why it can only be uncovered by analysis. The established tools range from the simple checklist to sophisticated modelling software:
- Five whys — by repeatedly asking “why?” you get from the symptom to the root cause. In detail: five whys.
- Ishikawa (fishbone) cause-and-effect diagram — arranges the possible causes into main categories (6M) (see below).
- Pareto analysis (80/20) — about 80% of the problems are caused by a few critical causes (~20%); it points you to the “critical few” (related to criticality analysis).
- Fault tree analysis — starting from the last failure, it traces the causes backwards until the trail can no longer be followed.
- Checklist — a structured data-collection form for the investigation.
This is how to link them together. Start with a fishbone to collect all root cause candidates; use a fact tree (fact tree analysis) to put the facts and the underlying causes into logical order, working backwards with the question “why did this happen?”; finally, the five whys closes and summarizes the analysis. The depth model is the same throughout: event → direct cause → basic causes → lack of control → loss. The toolbox is broader than this: barrier analysis, causal factor tree, cause map, change analysis, FMEA, and in the modern schemes the bow tie, the “events and causal factors” framework and Reality Charting.
The Ishikawa fishbone (6M)
Section titled “The Ishikawa fishbone (6M)”
Figure 3 — the Ishikawa diagram breaks the problem (the “fish head”) down into the main categories of possible causes (the “bones”).
The classic 6M categories walk through the whole space of possible causes: People, Machine, Method, Material, Measurement and Environment. The fishbone organizes the team’s ideas and prevents jumping to a cause too early; the candidates are then confirmed or ruled out by data and hypothesis analysis.
The types of RCA
Section titled “The types of RCA”Broadly speaking, RCAs fall into four application categories: safety-based (accident and incident investigation), production-based (quality and performance problems), process-based (faults in the process steps) and equipment-failure-based (RCFA, equipment breakdowns). Behind the schools stand common, universal principles: the root cause instead of the symptom, evidence-based work, and the prevention of recurrence.
RCA and FMEA — reactive and proactive
Section titled “RCA and FMEA — reactive and proactive”RCA and FMEA are two sides of the same coin: the failure mode that has already occurred is examined reactively by RCA (RCFA, Pareto, MTBF), while the one that has not yet occurred is analyzed proactively by FMEA.
Figure 4 — FMEA predicts and prevents before the failure occurs (proactive); RCA investigates and learns after the failure (reactive). The two methods complement each other; together they close the loop of prevention.
In a mature program the two are linked: the root causes uncovered in the RCA can be fed back into the FMEA (updated failure modes, causes, controls), while the high-risk failure modes of the FMEA mark out where it is worth preparing in advance. RCA resources are directed first at the “bad actors”: the equipment that fails repeatedly and with large losses, which the analysis of business losses identifies as critical equipment.
Industrial and safety context
Section titled “Industrial and safety context”The logic of RCA is industry-independent: from manufacturing through the energy sector to the process industry it is everywhere the same. Unfortunately, most organizations perform RCA only for reasons of compliance: when someone has been injured, or serious damage or environmental pollution has occurred, for which the authority also prescribes an investigation.
In a hazardous (Seveso) plant RCA is closely tied to incident and near-miss investigation (near miss): the root causes uncovered also strengthen the safety layers (re-checking the assumptions of hazop / LOPA). The real value lies in proactive use: the RCA of recurring, “small” equipment failures prevents the large events, and it feeds the reliability strategy and preventive maintenance.
Putting it into practice
Section titled “Putting it into practice”RCA comes alive when it becomes routine: trigger, team, evidence, verification. This is how to run it through an event:
- Rank the event by severity, and decide whether a quick report or a full investigation is needed.
- Convene the cross-functional team (operations, maintenance, technology, quality, HSE); at high severity, ask for an external, objective pair of eyes.
- Record the timeline and the sequence of events, then gather the evidence through interviews and workshops, working from the bottom up.
- Prove or rule out the root cause candidates with data and hypothesis analysis; never act on a presumed cause.
- Rank the root causes: a consequence score and a likelihood score (both 1–3), and the product (1–9) gives the order; develop countermeasures only for the highest ones. The scores are illustrative; the scale is set by the organization.
- Implement the actions with an owner and a deadline, then verify the result, and iterate if the problem comes back.
- Pass the lesson on to sister equipment and to the other sites: RCA is part of continuous improvement (kaizen), not a one-off report.
Common mistakes
Section titled “Common mistakes”- Stopping at the symptom. “Fixing” the surface failure brings the problem back. Instead: work through all six steps, all the way to verification.
- Jumping to a single cause. Focusing too early on one presumed cause. Instead: use a fishbone to collect all the candidates first, and only then narrow down.
- Looking for the “guilty person”. 60–80% of failures have a human dimension, but naming an individual as responsible masks the system-level cause. Instead: handle the human factor at the level of the missing procedure, the training, the supervision or the communication.
- Working without evidence. A “root cause” built on assumption. Instead: a documented timeline, and the naming of the candidates that were ruled out.
- Doing it only for compliance. RCA as a tick-box exercise, with no learning. Instead: use the severity-based ranking to select the important events, and put a full analysis on those.
- Skipping the verification. Instead: close the loop, and look at it again 12–18 months later.
Five typical reasons: the cause was never truly understood; an unexpected new combination of events or factors arose; the defences or the actions do not target the real cause; the action was not implemented; or it was implemented, but the system does not support its consistent use.
When NOT to use it (the limits of the method)
Section titled “When NOT to use it (the limits of the method)”RCA is a powerful tool, but it is not the right answer in every situation.
| Situation | Why not RCA | The right answer |
|---|---|---|
| A low-severity, one-off event | the full investigation costs more than it returns | a quick report from the asset team |
| The failure mode has not yet occurred | RCA is reactive; it works from an event that has happened | proactive FMEA and criticality analysis |
| You want to investigate every failure | resources spread thin deliver only mediocre work everywhere | a severity-based top list, a “bad actor” focus |
| A certified protection layer is needed | the lesson of an RCA does not become an audited barrier | design according to LOPA/SIL and hazop |
| There is no management commitment | a root cause without implementation is worth nothing | first an owner, a deadline, a follow-up cadence |
Take it home (keys)
Section titled “Take it home (keys)”- Rank first, analyze second. Severity decides who leads the investigation and what the mandatory output is.
- Keep the order: fishbone for the candidates, fact tree for the logic, five whys for the closing, down to the lack of control.
- Prove, don’t presume. Whatever does not rest on a timeline and on data is a hypothesis; write down the exclusions too.
- Close the loop twice: after implementation, and then 12–18 months later in a separate review.
- Spread the lesson to sister equipment, and feed it back into the FMEA.
Self-test
Section titled “Self-test”- “The bearing overheated”: is this a symptom, a problem or a root cause?
- For a short, low-loss, one-off shutdown, what is the proportionate investigation, and who leads it?
- Why is it not enough if the five whys stops at an operator error?
Answer key: 1) A symptom (the observable state caused by the problem). · 2) Low severity: a quick report at the asset team level is enough. · 3) Because you have to get from the direct cause to the basic causes and to the lack of control (procedure, training, supervision), otherwise the action does not target the real cause.
Hands-on
Section titled “Hands-on”- Run a five whys analysis on a recent failure, score the root causes you find on consequence × likelihood (1–3 × 1–3), and select the one with the highest score.
How does this show up in digital practice?
Section titled “How does this show up in digital practice?”RCA does not stop at the paper-based investigation report: the same principle is realized in software too. The mechanism differs, the principle is the same: recorded evidence, tracked action, measurable recurrence.
| RCA element | Digital implementation | What it delivers |
|---|---|---|
| Evidence and timeline | e-log, historian trend, event timestamp | facts instead of recollections |
| Recording the failure mode | CMMS work order with a structured failure code | an analyzable failure history, MTBF, Pareto |
| Early signal of the root cause | PdM / condition-based alert | the degradation shows up before the shutdown |
| Action tracking | a digital register with an owner and a deadline | the loop is demonstrably closed |
| “Has it come back?” | asset condition dashboard, recurrence report | the verification is measurable |
Modern digital systems realize the same principles: if the failure is recorded with evidence, a failure code and a tracked action, the analysis can work from data rather than from memory.
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”RCA connects to the shift log at two points. First, the trigger and the evidence: recurring abnormalities, near misses and small failures recorded in the shift log (OPEREX) give the timeline and sequence of events. Second, the closing of the loop: the corrective actions (what / who / by when) and the “has it come back?” check are auditably trackable here.
Terminology (HU / EN)
Section titled “Terminology (HU / EN)”| Hungarian | English (canonical) | Abbreviation |
|---|---|---|
| Gyökérok-elemzés | Root Cause Analysis | RCA |
| Gyökérok-meghibásodás elemzés | Root Cause Failure Analysis | RCFA |
| Gyökérok | Root cause | — |
| Tünet / szimptóma | Symptom | — |
| Okozati (közrejátszó) tényező | Causal / contributing factor | — |
| Közvetlen ok / alapok / a kontroll hiánya | Direct cause / basic cause / lack of control | — |
| Ok-okozati diagram (halszálka) | Cause-and-effect (Ishikawa / fishbone) diagram | — |
| 5 Miért | Five Whys | 5W |
| Hibafa-elemzés | Fault Tree Analysis | FTA |
| Tényfa-elemzés | Fact Tree Analysis | FTA |
| Pareto-elemzés | Pareto analysis (80/20) | — |
Note: the abbreviation “FTA” denotes two different tools. The fault tree breaks the unwanted event down into the combinations of contributing failures; the fact tree links the objective facts and the underlying causes back into logical order. In the investigation schemes it is typically the latter that appears.
Which failures should be investigated?
Not all of them. The level is decided by severity: at low severity a quick report from the asset team is enough, at medium an appointed investigator and a short summary are needed, and at high severity involving an external expert is also justified. The order comes from a top list maintained on the basis of total business and HSE impact.
How long should you keep asking "why"?
Until you get from the direct cause, through the basic causes, to the lack of control — that is, until you reach a cause the organization is able to eliminate. If the answer is merely the name of a person, you stopped too early.
What are the six steps of root cause analysis?
(1) define the problem, (2) collect data and evidence, (3) identify the possible causal factors, (4) develop solutions, (5) implement, (6) track the effectiveness — iteratively.
What tools does RCA use?
The five whys, the Ishikawa (fishbone) diagram, fact tree and fault tree analysis, Pareto analysis (80/20) and checklists; more broadly, barrier analysis, the cause map, change analysis and FMEA as well.
How does RCA differ from FMEA?
RCA is reactive (it investigates after the failure), while FMEA is proactive (it predicts and prevents before the failure). The two complement each other: the root causes from an RCA can be fed back into the FMEA.
Related concepts
Section titled “Related concepts”five whys · FMEA · criticality analysis · reliability strategy · preventive maintenance · A3 report · kaizen · near miss · hazop
Next step
Section titled “Next step”If you have understood this, from here it is worth going on — in this order:
- five whys — the closing tool in detail: from the direct cause to the basic causes and the lack of control.
- FMEA — the proactive mirror image: the analysis of failure modes that have not yet occurred, where you feed back the lesson of the RCA.
- reliability strategy — how root causes turn into a maintenance strategy.
References / further reading
Section titled “References / further reading”- Reliabilityweb.com — Uptime Elements: the reliability framework that treats root cause analysis as a fundamental element of asset management.
- ASQ (American Society for Quality) — Root Cause Analysis: a public methodological summary of the RCA tools.
- IEC 61025 — Fault Tree Analysis: the international standard for fault tree analysis.
- ISO 14224: reliability and maintenance data collection in the process industry; the reference basis for failure mode classification.
- Kaoru Ishikawa: Guide to Quality Control. Asian Productivity Organization, 1976 — the foundational work on the fishbone diagram.
In practice
Recurring abnormalities and near misses in the shift log (OPEREX) give both the trigger for an RCA and its evidence base (timeline, sequence of events); the corrective actions (who / by when) and the post-fix »has it come back?« check are auditably trackable in the same place, so the loop can be closed from the symptom to the verified elimination of the root cause.
Learn more: Maintenance →