TPM — Total Productive Maintenance
≈ 21 min read · 4,246 words
Your car has a service book, and you check the oil level and tyre pressure yourself — you don’t wait until it breaks down on the motorway. You clean and lubricate your household appliances and listen for unusual noises. This everyday “I look after it so it doesn’t break” idea is exactly what TPM (Total Productive Maintenance) does at plant scale: it does not leave the machine in the maintenance crew’s “no man’s land,” but also draws in the operator who runs it into the daily care, so that an emerging fault surfaces early — not as downtime or as an accident. Let’s look at what it is, why it works, and how it is measured.
TPM (Total Productive Maintenance) is a maintenance system that also involves the operators, in the service of equipment reliability. It treats maintenance not as the separate task of the maintenance department, but as a shared activity that spans the whole organization: its aim is to prevent failures, minimize downtime and equipment losses, and make machines safer and more economical. “Total” carries three meanings: the participation of all employees, the elimination of all sources of loss, and care that spans the equipment’s entire life cycle. The P in the name is Productive, not Preventive — the performance of TPM is measured with the OEE (Overall Equipment Effectiveness) metric.
Figure 1 — the TPM “house”: at the top the total productive (Productive) maintenance, standing on five pillars, with higher OEE as its foundation (availability × performance × quality).
Who is this for?
Section titled “Who is this for?”This article is for those who deal with equipment reliability and the maintenance culture in practice: operator · production and plant manager · shift supervisor · maintenance and reliability engineer · process engineer · Lean/CI specialist · plant economist · HSE.
Learning objectives
Section titled “Learning objectives”After reading this article you will be able to:
- explain why the P in the name TPM is Productive (not Preventive), and what the three meanings of “total” are;
- list the five pillars and the six equipment losses, and say which loss degrades which OEE factor;
- calculate an OEE from the three factors, and explain why the product nature is deceptive;
- recognize when TPM is not the primary answer (a design fault, a certified safety function);
- outline a gradual path for a TPM rollout (5s → autonomous maintenance → fact-based planned maintenance).
In brief (TL;DR)
Section titled “In brief (TL;DR)”- TPM (total productive maintenance) is a maintenance system that improves equipment availability and reliability and also involves the operators — not “preventive,” but Productive Maintenance.
- It is built on five pillars: (1) loss-reduction improvements, (2) autonomous maintenance (operator maintenance), (3) planned maintenance based on failure history, (4) training of operators and maintenance staff, (5) early equipment management against startup losses.
- The method fights six equipment losses: breakdown, changeover/setup, minor stoppages, speed loss, scrap/rework, and startup (ramp-up) loss.
- The performance metric is OEE = availability × performance × quality; because these three factors are multiplied, the aggregate effectiveness can drop seriously even when each part is individually good.
- Preventive maintenance is an important element of TPM: it is cheaper to prevent a failure than to fix it afterwards; a significant part of the routine is done in standardized form by the operators.
- Its origins reach back to Japanese preventive maintenance in the 1950s (Deming’s influence); the first full rollout was at DENSO (1960), and the method was consolidated by Seiichi Nakajima (TPM Nyumon, 1984).
- In the process industries availability is often the largest loss within OEE — and TPM works here too: plants producing liquid-phase products have won the TPM Prize (Japan Institute of Plant Maintenance).
Why it matters (the stakes)
Section titled “Why it matters (the stakes)”A minor equipment abnormality — slight vibration, a drop of a leak, an unusual noise — is cheap on its own. The trouble is the chain reaction it sets off if no one notices and no one intervenes: incipient wear turns into a breakdown, that into an unplanned stoppage, that into lost production and — in continuous operation — into a risky startup/shutdown transient. The same deviation can be solved with a shot of lubrication when it arises, but at the end of the chain it is orders of magnitude more expensive and more dangerous.
Figure 2 — the escalation of neglected equipment care: the further down the chain you catch it, the more expensive and riskier it is. TPM breaks it at the very start of the chain — at the operator’s daily inspection.
The lesson is simple: the cheapest failure is the one that never happens. That is why it pays to invest attention at the start of the chain — in daily care and fact-based prevention — rather than fixing it expensively at the end.
What is TPM, and where does it come from?
Section titled “What is TPM, and where does it come from?”TPM is a maintenance methodology aimed at improving productivity, which turns caring for equipment into a shared task that spans the whole organization. Its starting point is that in the traditional model — where running the machine belongs to the operator, and maintenance belongs solely to the maintenance department — the state of the equipment falls into “no man’s land”: the operator does not feel that the machine’s health is theirs, and the maintenance crew only sees it when there is already trouble. TPM tears down this wall.
The “total” qualifier means three things at once:
- the participation of all employees — from the operator through the maintenance technician to the engineer and management;
- the elimination of all sources of loss — not only the spectacular machine breakdown, but also the small stoppages and the speed loss;
- the equipment’s entire life cycle — from design and commissioning to decommissioning.
Origin and history. The roots of TPM reach back to the early 1950s, when — under the influence of Deming’s work — preventive maintenance began to be applied in Japan. The first company to introduce preventive maintenance in full was the Japanese firm DENSO, in 1960. With the rapid spread of automation the demand for maintenance work suddenly grew, so the usual maintenance activities had to be made performable by the operators too — this is where the idea of autonomous maintenance comes from. The foundations of TPM were consolidated by Seiichi Nakajima in his book TPM Nyumon (Introduction to TPM) in 1984; in 1987–88 Nakajima’s US lecture tour and the English edition of his book launched the spread of TPM in the USA, and a few years later in Europe.
In the name, the P = Productive, not Preventive. Preventive maintenance is one element of TPM, but it is not TPM itself. TPM is broader than that: it not only prevents failures but also improves the equipment, and makes maintenance a tool of productivity (not just of fault-fixing).
How does TPM work?
Section titled “How does TPM work?”TPM is built on five pillars, works against six equipment losses, and makes preventive maintenance part of the operator’s daily routine. Let’s take them in turn.
What are the five pillars?
Section titled “What are the five pillars?”TPM is usually described with five pillars:
- Loss-reduction (kaizen) improvements — targeted improvement activities against the six equipment losses (see below).
- Autonomous maintenance — many routine activities are done by the operator, not by the maintenance department. These operations (cleaning, lubrication, checking, minor adjustments) generally require no special qualification, only knowledge of the machine — but they are precisely defined and standardized. This way the operator “takes the machine in hand” every day and notices an emerging abnormality early.
- Planned maintenance — scheduling maintenance based on the failure history, not purely on a calendar/time basis. That is, not “every machine every X months,” but intervening where and how the actual data justify it.
- Training of operators and maintenance staff — continuous development of operating and maintenance skills; the operator can only perform autonomous maintenance if they understand how the machine works and the signs of an abnormality.
- Early equipment management — prevention built in already at design and commissioning, to avoid the losses that appear when new equipment is started up; the experience is fed back into the design of the next machine.
How does autonomous maintenance close the loop?
Section titled “How does autonomous maintenance close the loop?”The second pillar is the “engine” of TPM: the operator’s daily routine makes the machine’s state visible and produces the data that planned maintenance works from. If the operator cleans, lubricates and inspects every day, an emerging abnormality (vibration, leak, change in sound) surfaces early, is captured in the failure history, and planned maintenance is directed exactly where the actual data justify it — so there is no unexpected stoppage.
Figure 3 — the autonomous maintenance loop: operator’s daily routine → visible machine state → recording the abnormality → fact-based planned maintenance → no unexpected stoppage.
What are the six equipment losses?
Section titled “What are the six equipment losses?”TPM attacks six sources of loss. These are the classic “six big losses,” each of which degrades OEE:
| # | Loss (HU) | Loss (EN) | OEE factor it degrades |
|---|---|---|---|
| 1 | Meghibásodás (géptörés) | Breakdown loss | Availability |
| 2 | Átállás / beállítás | Setup & adjustment loss | Availability |
| 3 | Kisebb leállások (mikroleállás) | Minor stoppage loss | Performance |
| 4 | Sebességcsökkenés | Speed loss | Performance |
| 5 | Selejt és újramunkálás | Quality defects & rework | Quality |
| 6 | Indítási (felfutási) hozamveszteség | Startup yield loss | Quality / performance |
The point: losses 1–2 increase downtime (availability), 3–4 reduce the pace (performance), and 5–6 degrade the good yield (quality). TPM tackles all three fronts at once.
Why “Productive” and not “Preventive”?
Section titled “Why “Productive” and not “Preventive”?”Preventive maintenance is an important element of TPM. Its basic idea is that it is much better to prevent a failure than to solve the problem after it has occurred. Since the number of machines multiplied, a significant part of preventive maintenance is done by the operators — in the form of precisely defined, standardized operations. An important distinction: TPM’s planned maintenance is not timed maintenance, but a fact-based intervention built on the failure history. This is what links autonomous maintenance (where the data arises) with planned maintenance (where the data is used).
Process-industry context + safety
Section titled “Process-industry context + safety”In the process industries (oil refining, petrochemicals, chemicals) the logic of TPM applies strongly, but the emphasis shifts. Here the stakes are not the reliability of discrete-piece production lines but of continuously operating equipment (distillation columns, compressors, pumps, heat exchangers, reactors), and the key metric is mechanical availability — that is, whether the equipment is operable during the planned time.
- Availability is often the largest loss. When Lean is introduced, in most plants equipment availability is a significant — often the largest — source of process losses among OEE’s three factors. TPM is therefore a strong tool for improving overall performance in the process industries too.
- An unplanned stoppage is expensive and risky. In continuous operation an unexpected compressor or pump trip is not only lost production but also a startup/shutdown transient, which is among the riskiest operating states in the process industries. TPM’s autonomous inspection (early detection of vibration, sound, leak, temperature) thus carries direct process-safety value too.
- Safety connection. One explicit aim of TPM is to make machines safer. A well-maintained, leak-free, clean piece of equipment is less likely to lead to a release of hazardous material. In a Seveso-classified plant, up-to-date knowledge of the state of seals, valves and safety devices, and regular, standardized inspection, directly support the process safety goals.
- TPM is proven in the process industries too. Among the plants operating in the liquid-phase products industry, Japanese butyl plants were the first to win the TPM Prize (Total Productive Maintenance award) from the Japan Institute of Plant Maintenance — proving that the method applies not only to discrete manufacturing but also to continuous process industries.
TPM improves the reliability and safety of equipment, but does not replace certified functional safety systems. The emergency-shutdown and protection logic must always be designed according to risk analysis and the relevant standards (LOPA, SIL / IEC 61511); equipment modifications are steered by moc (Management of Change) so that they do not introduce new risk. Give a specific numerical safety requirement only on the basis of the text of the relevant regulation/standard.
Putting it into practice (roadmap)
Section titled “Putting it into practice (roadmap)”A TPM rollout is typically built on 5s and on a stable, measured base. One possible, gradual path — in action-first steps:
- Lay the base with 5s and measurement. Without a clean, orderly work area the abnormality is not visible. In parallel, introduce downtime measurement — in the process industries most plants initially do not know their actual availability.
- Choose a pilot piece of equipment. Pick a bottleneck or critical piece of equipment where the reliability improvement has a direct business impact (in the process industries typically the capacity-determining unit).
- Assess the baseline OEE. Calculate the current OEE (see the measurement section) and break it down into the six losses — this shows which pillar brings the most.
- Start autonomous maintenance. An initial deep clean, then standardized daily clean–inspect–lubricate (CIL) for the operators; visual standards, check points on the machine.
- Collect the failure history. Record every stoppage cause, and from the data build up the fact-based (not calendar-based) planned maintenance.
- Train the operators and maintenance staff to recognize and handle emerging faults on the spot.
- Feed the lessons back into the design of new investments (early equipment management), so that startup losses are already reduced at commissioning.
- Standardize and spread the pilot results, then take it out to the next piece of equipment.
Experience shows that many companies achieve the improvement but few keep it. Without standardization and institutionalization (sustaining), the system slips back — plan for this, don’t retrofit it later.
A practical mini-scenario
Section titled “A practical mini-scenario”How would you start on a pump tomorrow? Pick a painful, frequently failing pump. Write a half-page daily inspection card for it: vibration by hand/hearing, bearing temperature, gland-seal leak, oil level, unusual noise — each with an “OK / flag” mark. The operator fills it in every shift and records the deviation immediately into the failure history. After two or three shifts you will see which sign keeps coming back — that gives you the first fact-based planned-maintenance item. A single card, zero investment, and yet the TPM loop works.
How do we measure TPM? (OEE)
Section titled “How do we measure TPM? (OEE)”The primary metric of TPM is the OEE (Overall Equipment Effectiveness / Efficiency): it compresses into a single number how well, how fast and how reliably the equipment produces. OEE is the product of three operational parameters:
OEE = Availability × Performance × Quality [%]
The three factors:
- Availability = (production time − downtime) / production time That is, how much of the planned operating time the machine actually produced (breakdown, changeover subtracted).
- Performance = (produced quantity × cycle time) / (production time − downtime) That is, how much it achieved during the actual run compared with the nominal pace (minor stoppage, speed loss subtracted).
- Quality = (produced quantity − scrap) / produced quantity That is, how much of what was made is good, sellable product.
The trap of the product nature. Because the three factors are multiplied, even individually “good-looking” values give a weak OEE: if all three factors are 95%, the OEE is only ≈ 85.7% (0.95 × 0.95 × 0.95). This is why it is not enough to focus on one factor at a time — all three must be improved together.
Worked example (from the source, illustrative). Take the following data for a line:
- Planned production time: 20.5 hours (from 24 hours, subtracting per shift 1 hour of break and 0.5 hour of planned preventive maintenance).
- Unplanned downtime: 1.5 hours.
- Planned cycle time: 30 s/pc.
- Actual production: 2020 pcs, of which 50 scrap → 1970 sellable pcs.
Then:
| Factor | Calculation | Value |
|---|---|---|
| Availability | (20.5 − 1.5) / 20.5 | 0.927 |
| Quality | 1970 / 2020 | 0.975 |
| Performance | 2020 / [(20.5 − 1.5) × (3600/30)] | 0.886 |
| OEE | 0.927 × 0.975 × 0.886 | ≈ 0.801 (≈ 80%) |
The breakdown obtained this way tells you where to allocate the resource: in this example about 2.5% quality, 7.3% availability, and the remaining performance (cycle-time) loss. OEE is thus a prioritizing tool: it shows whether quality, availability or cycle time needs improving first.
Audit points. In a TPM audit, look at: whether autonomous maintenance is standardized (is there a CIL standard on the machine, is it followed); whether the failure history is actually kept and used for scheduling planned maintenance; whether real, measured OEE/availability data is available (not an estimate); and whether the improvements are sustained (sustaining).
Common mistakes
Section titled “Common mistakes”The pitfalls of TPM almost all stem from the same thing: one element of the method is mistaken for the whole. In anti-pattern ↔ correction pairs:
- Conflating “Preventive” and “Productive.” Stopping at timed preventive maintenance, without autonomous maintenance and operator involvement — this is not TPM. Instead: bring the operator into the daily care, and make planned maintenance fact-based.
- Calendar-based maintenance instead of fact-based. “Every machine every X months” wastes when over-maintaining and leads to failures when under-maintaining. Instead: build planned maintenance on the failure history, not on the calendar.
- The lack of downtime measurement. If availability is not measured, OEE is unknown and prioritization happens blindly; the typical reflex (overtime, “pushing”) is only symptomatic treatment. Instead: first measure downtime and OEE, only then intervene.
- Leaving out the operator. If autonomous maintenance is not standardized and operators are not trained for it, TPM remains “extra work for the maintenance crew,” and no early signal is produced. Instead: a standardized CIL card + training, so the operator sees the abnormality.
- Optimizing the OEE factors separately. Because of the product nature, “sharpening” one factor at the expense of the other two can easily worsen the aggregate OEE. Instead: prioritize from the breakdown by the six losses, and improve all three factors together.
- Failing to sustain. After the improvement is achieved, standardization is skipped and the system slips back — this is the most common mistake. Instead: plan standardization and rollout as part of the introduction, not as an afterthought.
When NOT to use it (the limits)
Section titled “When NOT to use it (the limits)”TPM is strong but not universal. Knowing where TPM is not the primary answer is just as important as the method itself:
| Situation | Why (primarily) not TPM | The right answer | |
|---|---|---|---|
| A design / fitness fault in the machine | not even the best maintenance makes an under-designed machine reliable | redesign, root cause in the design ([[fmea.en | fmea]], early equipment management) |
| A certified safety function is needed (emergency shutdown, interlock) | TPM is a management/maintenance principle, not a certified protection layer | design per [[lopa-sil.en | SIL/LOPA]], IEC 61511 |
| A one-off, non-recurring failure | there is no durable pattern worth building into the daily routine | one-off root-cause analysis ([[5-miert.en | 5 Whys]]), recording the lesson |
| No measured base (no [[5s.en | 5s]], no downtime data) | TPM’s “abnormality” principle runs empty without a reference point | first stabilize: 5S + downtime measurement, then TPM |
Rule of thumb: TPM is strongest for recurring, measurable equipment losses. It does not replace a design fault, certified safety, or a missing measured base — it complements or is a precondition for them.
Key takeaways
Section titled “Key takeaways”- P = Productive, not Preventive — preventive maintenance is only one element; TPM also involves the operator and improves the equipment.
- The operator is the machine’s first “sensor” — the daily clean–lubricate–inspect catches the emerging fault before it becomes a stoppage.
- Fact-based, not calendar-based planned maintenance: the failure history tells you where and when to intervene.
- OEE prioritizes: the availability × performance × quality breakdown shows which loss to put resource on first.
- The product is deceptive: three times 95% is still only ~86% OEE — improve all three factors together.
- Sustaining is the hardest — standardize and spread the improvement, otherwise it slips back.
Self-check
Section titled “Self-check”- Why is the P in the name TPM “Productive,” and what is the difference between TPM and plain preventive maintenance?
- A machine’s OEE is weak. How would you use the six losses and the three OEE factors to find out whether availability, pace or quality needs improving first?
- Where is the boundary between TPM’s autonomous inspection and a certified emergency shutdown (SIL/ESD) — why does neither replace the other?
How does this show up in digital practice?
Section titled “How does this show up in digital practice?”The TPM principle does not end in the machine hall: the same logic is realized in software too. Instead of the paper CIL card and the wall-mounted OEE board, here a mobile operator round, automatic downtime collection and condition-based alerting carry the “care for it → measure → prevent” triad — the mechanism differs, the principle is the same.
| TPM principle | Digital implementation | What it delivers |
|---|---|---|
| Autonomous maintenance (daily CIL) | mobile operator round / digital check card, with enforced steps | the daily routine is completed and recorded auditably |
| Failure history | deviation log, event-linked recording, searchable machine history | the factual base for planned maintenance |
| Planned maintenance | CMMS work order based on fault history / condition | the intervention goes where the actual data justify it |
| Six losses + OEE | automatic downtime collection, OEE calculation for the three factors | a real, not estimated performance picture, prioritized |
| Early signal | condition-based (PdM) alert on vibration, temperature | the emerging abnormality surfaces before a stoppage |
| Standard work | a digital template captures the correct inspection sequence | no ad-hoc execution that differs from person to person |
Lean is not made only of shop-floor tools: modern digital systems realize the same principles in software. If a system makes the operator the machine’s first sensor and builds fact-based maintenance from the daily signals, chances are the logic of TPM is at work in the background.
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”TPM is a data-hungry method: both autonomous maintenance and planned maintenance only work if daily, shift-level data is produced about the equipment. The OPEREX shift log covers exactly this layer: shift by shift you can record and retrieve (1) the daily cleaning, lubrication and inspection of autonomous maintenance, (2) the abnormality detected by the operator (vibration, leak, change in sound or temperature), (3) the cause of stoppages by the six losses, and (4) the downtime data from which the availability factor of OEE can be calculated. The failure history produced this way gives the factual base for scheduling planned maintenance, and the per-shift trail is auditable evidence of compliance with the maintenance standards.
Terminology (HU / EN / JP — where relevant)
Section titled “Terminology (HU / EN / JP — where relevant)”| Hungarian | English | 日本語 (romaji) | Note |
|---|---|---|---|
| teljes körű hatékony karbantartás | Total Productive Maintenance (TPM) | 生産保全 (seisan hozen) | P = Productive, not Preventive |
| autonóm / önálló karbantartás | autonomous maintenance | 自主保全 (jishu hozen) | the operator does the routine |
| tervezett karbantartás | planned maintenance | 計画保全 (keikaku hozen) | based on failure history |
| megelőző karbantartás | preventive maintenance | 予防保全 (yobo hozen) | one element of TPM |
| korai berendezés-menedzsment | early equipment management | — | against startup losses |
| berendezés-hatékonyság (mutató) | Overall Equipment Effectiveness (OEE) | — | A × P × Q |
| rendelkezésre állás | availability | — | OEE factor |
| teljesítmény | performance | — | OEE factor |
| minőség(i hozam) | quality (yield) | — | OEE factor |
| hat (nagy) veszteség | six (big) losses | — | the targets of TPM |
| állásidő | downtime | — | planned or unplanned |
What is the difference between TPM and preventive maintenance?
Preventive maintenance is an activity aimed at preventing failures, and it is one element of TPM. TPM is broader than that: it involves all employees (autonomous maintenance), attacks each of the six losses, builds planned maintenance on the failure history, and spans the equipment’s entire life cycle. In the name, the P is “Productive,” not “Preventive.”
Why is OEE important in TPM?
OEE is TPM’s primary performance metric, compressing availability, performance and quality into a single number. Because it is the product of the three factors, it shows which type of loss drags overall performance down the most, and thus prioritizes where it is worth putting resource first.
What is autonomous maintenance, and why does the operator do it?
Autonomous maintenance is the handover of the daily routine (cleaning, lubrication, inspection, minor adjustment) to the operator. These operations require no special qualification, only knowledge of the machine, but they are standardized. Its advantage is that the operator “takes the machine in hand” every day and notices an emerging abnormality early — before it grows into a stoppage.
Does TPM work in the process industries (oil refining, chemicals)?
Yes. In the process industries it is often precisely equipment availability that is the largest OEE loss, so TPM is a strong tool. As proof, industrial plants producing liquid-phase products were the first to win the TPM Prize from the Japan Institute of Plant Maintenance.
Where should I start a TPM rollout?
On a stable 5s base and on downtime measurement — in the process industries many plants initially do not even know their actual availability. From there: a pilot piece of equipment, a baseline OEE assessment, autonomous maintenance, collecting the failure history, then fact-based planned maintenance.
Related concepts
Section titled “Related concepts”- oee — TPM’s performance metric (availability × performance × quality)
- 5s — TPM’s stabilizing base; order and cleanliness make the abnormality visible
- smed — changeover-time reduction, directly against the 2nd loss (setup)
- muda — the seven wastes; TPM targets the equipment-side losses
- jidoka — defect detection by the machine, related to reducing quality losses
- poka-yoke — fail-safe solution, the tool-level counterpart of preventing equipment faults
- moc — Management of Change; safe handling of equipment modification
- standard work — the standardized operations of autonomous and preventive maintenance
Next step
Section titled “Next step”If you have understood this, from here it is worth going on — in this order:
- oee — TPM’s metric in detail: how to calculate, break down and prioritize the three factors. Start with this, because you measure TPM’s result on it.
- 5s — TPM’s stabilizing base: without order and cleanliness the abnormality is not visible. The rollout starts here.
- smed — targeted reduction of the changeover (setup) loss, directly on the 2nd of the six losses.
References / further reading
Section titled “References / further reading”- Seiichi Nakajima: Introduction to TPM: Total Productive Maintenance (originally TPM Nyumon, 1984; English edition: Productivity Press, 1988). — the canonical work consolidating the basic concepts of TPM.
- Lonnie Wilson: How to Implement Lean Manufacturing. McGraw-Hill. — a practical, worked-example presentation of OEE and the five pillars of TPM.
- Raymond C. Floyd: Liquid Lean: Developing Lean Culture in the Process Industries. CRC Press, 2010. — the application of TPM and OEE in the process industries, with the example of plants that won the TPM Prize.
In practice
The daily cleaning, lubrication and inspection steps of TPM's autonomous maintenance, the abnormalities the operator detects (vibration, leak, change in sound), and the per-shift downtime causes of the six losses can be logged and retrieved shift by shift in the OPEREX shift diary, so that credible downtime data is produced for the OEE calculation and a failure history for preventive maintenance.
Learn more: Shift log →