Skip to content

Reliability strategy and RCM

≈ 17 min read · 3,392 words

In your car you run the alternator to failure, you change the air filter by the odometer, and you check the tyres by eye. Three components, three different maintenance decisions; in a plant the same question waits on tens of thousands of failure modes, and there it is no longer gut feel that decides, but method.

Reliability strategy development (RSD) is the method for building a technically correct, cost-effective maintenance program for an asset.

Where there is no maintenance requirement yet, RSD develops one; where one already exists, it optimizes it. It rests on three proven techniques (reliability-centered maintenance, RCM; preventive maintenance optimization, PMO; failure mode and effects analysis, FMEA), and its end result is a combination of time-directed (TD), condition-directed (CD) and failure-finding (FF) tasks that keeps the assets operable in the current operating context. Its key principle is that every maintenance action is tied to a specific failure mode: this way the program rests on a documented technical basis, not on tasks done out of habit.

rsd-3technika-en.svg Figure 1 — the three techniques of the reliability strategy (RCM, PMO, FMEA) come together into one maintenance program built from TD + CD + FF tasks.

For those who plan maintenance work, carry it out or pay for it: reliability engineer · maintenance planner · plant manager · process engineer · maintenance craftsman · operator · HSE / process safety specialist.

After completing this module you will be able to:

  • name the three techniques (RCM, PMO, FMEA) and the four tactics (CD, TD, FF, RTF)
  • apply the seven questions of RCM to a failure mode
  • explain the P–F interval, and why random failures are handled by a condition-directed task
  • justify why every maintenance task is tied to a failure mode
  • The goal: the right maintenance, on the right asset, at the right time; neither too much, nor too little.
  • Three techniques: RCM (from function → to failure mode → to task), PMO (review of the existing PM), FMEA (taking stock of the failure modes).
  • Four tactics: condition-directed (CD), time-directed (TD), failure-finding (FF) and the deliberate run-to-failure (RTF).
  • Every task is tied to a failure mode: “we do it because we can” is not a good enough reason.
  • 77–92% of failures are random (not age-related) — these are handled only by a condition-directed (CD) task, not by periodic replacement.
  • RCM is a seven-question process to the SAE JA1011 standard; PMO is practically its reverse (it starts from the existing tasks), and it also adds new tasks for the uncovered failure modes.
  • Criticality drives the depth, because reactive maintenance typically costs 2–4× as much as planned work.

The end result of RSD (whether we build a new program or optimize an existing one) is a technically correct and cost-effective set of tasks that improves the reliability and operational availability of the system, and gives a documented technical basis for every maintenance decision. The three techniques are differently structured processes for the same goal, and each of them creates time-directed (TD), condition-directed (CD) and failure-finding (FF) tasks: together these make up the preventive maintenance program.

RCM is a systematic, disciplined process for creating a suitable maintenance plan for an asset/system that minimizes the probability of failures, ensuring safety, the operation of the system and conformance to its purpose. The concept started in the aircraft industry in the 1960s: the aviation authority showed that the amount of preventive maintenance and reliability are not proportional, and that some items cannot be maintained at all. In 1974 the US Department of Defense commissioned the report titled Reliability-Centered Maintenance (1978), and in the eighties the nuclear power industry took it over. Although its greatest benefit in the design and development phase lies in eliminating failure modes, it can be applied successfully at any point of the life cycle.

PMO — preventive maintenance optimization

Section titled “PMO — preventive maintenance optimization”

PMO evaluates every existing PM task and eliminates the unnecessary, redundant or useless activities, so that the scarce maintenance resource can be reallocated to the tasks that genuinely prevent failures. But PMO does not only delete: in most cases it also adds further tasks for the failure modes that the existing program does not cover at all. Its direction, however, is the reverse of RCM: it starts from the task and works backwards to the failure mode (is it relevant, is it effective?), while RCM starts from the function and the functional failure, and assigns a task forward to the failure mode.

The proven moves of PMO: replacing a calendar-based task with a running-hour-based, condition-directed or run-to-failure task where that is feasible; eliminating duplicated PM (when several groups do the same thing on the same asset); the correct split of tasks between maintenance and operations; improving the incomplete, sloppy or not cost-effective tasks; and keeping the program alive with regular updates.

Does this PM add value? An existing task stays in the program if all six statements are true of it; if even one is false, the task can be deleted.

  1. The failure is not detectable by the operator in normal operation.
  2. The inspection tells when the given degradation mechanism will cause a failure.
  3. If it is a PM task, it targets a clear wear-out mechanism.
  4. The inspection does not accelerate the degradation and does not introduce a new failure risk.
  5. There is no other protection or means of detection.
  6. Run-to-failure is not acceptable from a business point of view.

FMEA — failure mode and effects analysis

Section titled “FMEA — failure mode and effects analysis”

FMEA is the primary tool of the RCM analysis: it ensures that every failure mode is taken into account. (For the detail, see the FMEA article.) The quality management standard ISO/TS 16949, for instance, requires FMEA for the product, the design and the process alike.

  • Time-directed (TD): a task aimed at preventing the failure, based on calendar or operating time (periodic replacement, overhaul).
  • Condition-directed (CD): condition monitoring aimed at the onset of the failure symptom; see asset condition management.
  • Failure-finding (FF): a scheduled check of whether a hidden failure (not visible in normal operation) has occurred; typically on standby and safety systems.
  • Run-to-failure (RTF): a deliberate economic decision, when the effect and the cost of the failure are smaller than preventing it.

The process that answers the following seven essential questions is what we call reliability-centered maintenance. The minimum requirements are set out in the SAE JA1011 standard.

rcm-7kerdes-en.svg Figure 2 — the seven questions of RCM: from the functions to the appropriate proactive, or default, actions.

RCM is defined by four principles, which distinguish it from every other PM planning process:

  1. The primary goal is to preserve the function of the system — not the equipment itself, but what it delivers.
  2. The failure modes that block the function must be identified.
  3. The functional needs (failure modes) must be ranked — according to the consequences.
  4. Relevant and effective tasks must be chosen — carrying out a task is not justified in itself just because “we can” or “it belongs to the area”.

The right tactic depends on the nature of the failure mode. The proven decision order: look at the condition-directed (CD) task first, because it involves the least intervention, it is cheaper and faster, and it lets you plan the repair before the actual failure.

rcm-taktika-en.svg Figure 3 — tactic selection logic: condition-directed → time-directed → failure-finding → run-to-failure, and as a last resort, redesign.

  • A time-directed (TD) task is typically justified on a simple, single-piece item, where the age–reliability relationship is direct (metal fatigue, mechanical wear, a part designed as a consumable); here the age limit improves the reliability of the complex item it is part of.
  • A condition-directed (CD) task suits initial, random and “infant mortality” pattern failures, typically on complex items; its applicability is limited by the length of the P–F interval.
  • Failure-finding (FF) where the loss of function is not evident: it reduces the risk of the multiple (hidden + evident) failure to an acceptable level.
  • Run-to-failure (RTF), if the consequence is bearable and prevention is more expensive.
  • Redesign, if the consequence is severe and there is no good proactive task; this is the default output of the logic tree on a critical item.

The everyday anchor of the four tactics (a car):

Tactic Example Why this one
RTF alternator not critical, random failure, no measurable degradation mechanism
TD air filter cheap part, predictable life; with run-to-failure it would become a chronic problem
TD timing belt critical consequence, hard to measure, but predictable life
CD tyre, catalyst degradation measurable cost-effectively
Redesign any of them it is decided in the design phase

⚠️ A periodic overhaul can increase the total failure rate, because it introduces “infant mortality” into an otherwise stable system. This is why periodic replacement is not a good answer to every failure.

One of the most important insights of maintenance strategy is about the nature of failures. According to three major studies (the 1978 RCM report, Sweden 1973, US Navy 1983), 77–92% of failures are random (not age-related), and only 8–23% are age-related. (The classic bathtub curve and the six patterns in detail: the bathtub curve and failure patterns.)

megbizhatosag-mintazatok-en.svg Figure 4 — most failures are random (top); the P–F curve (bottom) shows when the failure becomes detectable (P) and when it turns into a functional failure (F).

This has two practical consequences:

  • Random failures can only be handled by a condition-directed (CD) task — periodic replacement/overhaul does not help against them (and may even harm).
  • The P–F interval (the time between the potential failure becoming detectable and the functional failure) determines whether a CD task can work at all, and how often it has to be done: this is the window in which, having detected the degradation, you can still intervene in time.

A practical rule of thumb for the frequency. Schedule the condition-directed inspection at a frequency corresponding to 0.5–1 times the P–F interval. For a time-directed task the interval is typically ~70% of the MTBF; if there is no MTBF data, start from 50% of the estimated value, then ask the question: what business risk would extending it carry? If that is acceptable, stretch the interval.

  • The cost of reactive maintenance falls, as failures are prevented and condition monitoring replaces the superfluous preventive tasks. Corrective reactive work, because of the low efficiency inherent in it, typically costs two to four times as much as planned work.
  • Cost-effectiveness is built in: maintenance at the right level according to criticality. Maintenance that is not cost-effective is identified but not carried out by the strategy.
  • PMO benefit: most existing PM programs cannot be traced back to their origin; PMO weeds out the superfluous and ineffective tasks in a structured way.
  • Documentation: the RCM/PMO/FMEA analysis puts the failure modes and the justification of the tasks in writing; this is good training material for new operators and maintenance craftsmen.
  • Condition-based, not calendar-based replacement: it extends the life of the equipment and the facility.
  • But it can be more expensive in the short term: because of buying the technology, the training and the initial condition survey, the program can increase the maintenance cost at first. This extra is relatively short-lived, but you need to know about it for the decision.

The logic is industry-independent: a critical asset is one that we classify as critical because of the effect its failure has on safety, the environment, quality, production and maintenance. In a hazardous (Seveso) plant the strategy is especially important at two points: the failure-finding (FF) tasks of the hidden functions (e.g. standby, shutdown and safety systems) make sure that the protection layer really works when it is needed (SIL/ESD context); and run-to-failure (RTF) is permitted only where the consequence is provably bearable, never on a safety or environmental failure mode. In RCM safety comes before economics: the cost of maintaining safe working conditions is not weighed as an RCM cost.

The application of RCM, PMO and FMEA varies, but most procedures contain several or all of the following nine steps:

  1. Select the system and gather the information relevant to it (the criticality marks out where the detailed analysis should go).
  2. Draw the system boundaries.
  3. Describe the system and make a functional block diagram.
  4. Take stock of the system functions and the functional failures.
  5. Run an FMEA with a cross-functional team: this is the primary tool for taking stock of every failure mode.
  6. Run the logic (decision) tree analysis (LTA) according to the consequences.
  7. Select the maintenance tasks failure mode by failure mode with the logic of Figure 3, and tie the task to the failure mode (this is the documented technical basis).
  8. Arrange the tasks into packages and implement them.
  9. Keep the program alive: continuous improvement from root cause analysis and from operational feedback.

With which tool? RCM has a whole family, with different strengths. Full RCM is thorough but “heavy”, and focuses purely on preserving the system function (less on the size of the business impact). RBI (risk-based inspection) is tailored and inspection-focused, with a menu of degradation mechanisms; it is a good way to define the turnaround work scope, less so for rotating equipment. RCM Lite / PM screening is very fast and effective, but it is only a first cut, not an optimization, and its error potential is greater depending on the analyst. Never accept the first pass: audit and coach the team.

  • Treating every failure as age-related — random failures cannot be handled by periodic replacement.
  • Too much / ineffective PM — a program swelling without any thought about the cost and the value of the tasks.
  • Periodic overhaul where it does harm — it introduces “infant mortality” into a stable system.
  • The task is not tied to a failure mode — there is no technical justification, only habit.
  • RTF on a safety failure mode — run-to-failure is permitted only where the consequence is bearable.
  • Ignoring hidden failures — without an FF task the standby/safety system can be “dead” unnoticed.

When NOT to use it? (the limits of the method)

Section titled “When NOT to use it? (the limits of the method)”
  • It does not fix a bad design. A maintenance program can only sustain the reliability that is built into the design of the system; no maintenance eliminates faulty design. This is why maintenance experience has to be fed back to the designers: see design for reliability.
  • It cannot be run on everything. Doing an FMEA properly is time-consuming and resource-hungry: attempting it on every asset would tie up the entire specialized engineering capacity, and the benefit would appear years later. This is why criticality steers it.
  • The intervention itself is a risk. Disassembly can introduce a fault into a working machine; hence the rule “if it isn’t broken, don’t fix it”.
  • An inspection showing the momentary state is not a prediction. An indicator lamp test tells you whether it is good now, but not when it will fail; on a protection and standby system, however, that is exactly the point of the failure-finding (FF) task.

The output of the reliability strategy is operator-friendly: the condition-directed (CD) and failure-finding (FF) tasks are typically round checks and measurements. A digital shift log (OPEREX) treats these as items that can be ticked off shift by shift, trends the condition parameters, and records the completion of the hidden-function (FF) checks in an auditable way. This way the maintenance program does not exist only on paper: shift after shift it is visible that the CD/FF tasks were really done, and the deteriorating trends (P–F interval) are escalated in time.

Hungarian English (canonical) Abbreviation
Megbízhatósági stratégia kidolgozása Reliability Strategy Development RSD
Megbízhatóság-központú karbantartás Reliability-Centered Maintenance RCM
Megelőző karbantartás optimalizálása Preventive Maintenance Optimization PMO
Idő-alapú feladat Time-Directed task TD
Állapotfüggő feladat Condition-Directed task CD
Hibakereső feladat Failure-Finding task FF
Hiba-kivárás Run-to-Failure RTF
Rejtett hiba Hidden failure
Üzemelési kontextus Operating context
P–F intervallum P–F interval
Kockázat-alapú ellenőrzés Risk-Based Inspection RBI

The terminology follows the usage of Uptime Elements and SAE JA1011.

What is reliability strategy (RSD)?

A systematic approach to developing a technically correct, cost-effective maintenance program for an asset, or to optimizing an existing one. It rests on three techniques (RCM, PMO, FMEA), and ties every task to a specific failure mode.

What is RCM, and which standard describes it?

Reliability-centered maintenance (RCM) is a disciplined process that creates a maintenance plan preserving the system function, with tasks assigned to failure modes. Its minimum requirements are set out in the SAE JA1011 standard, and it is based on answering seven questions.

What is the difference between RCM and PMO?

RCM starts from the functions and functional failures, and assigns tasks forward to the failure modes; PMO starts from the existing tasks and checks their relevance backwards. PMO is practically the reverse of RCM, and it is ideal for quickly optimizing existing PM programs. Importantly, PMO does not only weed out: in most cases it also adds new tasks for the failure modes the program has not covered so far.

Why doesn't periodic replacement solve everything?

Because 77–92% of failures are random, not age-related — these are handled only by condition-directed (CD) tasks. A periodic overhaul can, on top of that, introduce “infant” failures into a stable system, increasing the failure rate.

What is the P–F interval?

The time between the potential failure becoming detectable (P) and the functional failure (F). This decides whether a condition-directed task can work; schedule the inspection at 0.5–1 times the P–F interval.

When is run-to-failure (RTF) justified?

When the consequence and the cost of the failure are provably smaller than the cost of prevention — as a deliberate economic decision. Never on a failure mode with a safety or environmental consequence.

  • Every task is tied to a failure mode: this is the documented technical basis; “out of habit” is not a justification.
  • CD first, then TD: the condition-directed task solves more with less intervention, more cheaply and more predictably.
  • Frequency is not gut feel but a number: 0.5–1 × P–F, and for TD, ~70% of the MTBF.
  • Run-to-failure (RTF) is a decision, not an omission, but never with a safety or environmental consequence.
  • The hidden function has to be looked for: without an FF task the standby and safety system can be dead unnoticed.
  • The program is a living system: if it is not refreshed from failure experience, it swells back to the old task list.
  1. 77–92% of failures are random: which tactic dominates for these, and why is it not periodic replacement?
  2. On a CD inspection the P–F interval is 8 weeks. At what frequency do you schedule it, and on what rule?
  3. When is deliberate run-to-failure (RTF) justified, and when is it forbidden?

Answer key: 1) Condition-directed (CD); a time-based replacement does not protect against an age-independent failure, and may even introduce an “infant” fault. · 2) Every 4–8 weeks, that is, 0.5–1 times the P–F interval, so that you look at it at least once inside the window between P and F. · 3) If the consequence and the cost of the failure are provably smaller than preventing it; never on a failure mode with a safety or environmental consequence.

Applied task: choose a tactic (CD/TD/FF/RTF) for three of your own failure modes with the logic of Figure 3, then run the six value-check statements on their existing PMs.

  • SAE JA1011Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes: the seven questions and the minimum requirements of RCM. Its application guide is SAE JA1012.
  • F. S. Nowlan – H. F. Heap: Reliability-Centered Maintenance. US Department of Defense, 1978 — the original source of the six failure patterns.
  • John Moubray: Reliability-Centred Maintenance (RCM II). Butterworth-Heinemann, 1997.
  • Uptime Elements (Reliabilityweb.com) — the RSD element: the competence framework of RCM, PMO and FMEA.
  • ISO 14224 — the reference basis for the failure mode and consequence categories.
  • API RP 580 / 581 — risk-based inspection (RBI), the inspection-focused member of the RCM family.

FMEA | criticality analysis | preventive maintenance | asset condition management | root cause analysis | design for reliability | the bathtub curve | reliability KPIs | TPM | autonomous maintenance | risk management