Vibration analysis
≈ 13 min read · 2,638 words
On the spin cycle the washing machine sets off across the bathroom, and the desk fan rattles the board underneath it. Both are saying the same thing: something has gone out of balance, and the machine says so with vibration. On the bearing housing of a pump the same signal is measurable, only quieter, and long before the breakdown.
Vibration analysis detects developing machine faults by measuring and analyzing the vibration of rotating equipment, long before the breakdown.
Introduced and applied correctly, it is one of the five primary technologies of condition-based maintenance (CBM). It picks bearing faults, imbalance, misalignment, gearbox and electrical faults out of the spectrum or the waveform, so there is still time to plan the repair. The program is effective when it works with the data, the range and the alarm chosen for the failure modes of the FMEA.
Figure 1 — vibration analysis “sees” the typical faults of rotating machines: bearing, imbalance, misalignment, gearbox, electrical, looseness.
Who is this for?
Section titled “Who is this for?”This article is for those who decide on the condition of rotating machines: reliability engineer · maintenance engineer · mechanical engineer · vibration analyst · plant manager · shift supervisor · operator · process engineer.
Learning objectives
Section titled “Learning objectives”After reading this article you will be able to:
- name which rotating machine faults vibration analysis detects;
- select the data type, the Fmax and the resolution that fit the failure mode;
- distinguish an overall alarm from a focused alarm;
- derive the collection interval from the P–F interval;
- draw the line between a machine protection trip and a safety function.
In brief
Section titled “In brief”- Vibration analysis is the established CBM technique for the early fault detection of rotating machines, and the acceptance and quality-control tool of precision installation.
- Faults detected: bearing fault, imbalance, misalignment, gearbox and electrical faults, looseness.
- FMEA-driven: “vibration analysis” is a general statement; the failure mode is identified by targeted data.
- The rate of change matters just as much as the momentary value: the trend signals the development.
- The overall alarm is for tripping the machine; early detection calls for a focused alarm backed by data.
- Frequency: half the P–F interval, tightened when a fault is detected.
Why it matters (the stakes)
Section titled “Why it matters (the stakes)”The stake is whether the rotating machine stops on schedule or on its own. An unplanned shutdown brings secondary damage and lost production on top of the cost of the repair. Early detection turns this into planned work.
It is forbidden to use general data collection points on a “maybe we’ll get lucky” basis, and it is forbidden to build the program on a fixed, time-based interval without taking the failure modes into account.
What is vibration analysis, and which rotating machine faults does it detect?
Section titled “What is vibration analysis, and which rotating machine faults does it detect?”Vibration analysis is a widely used predictive maintenance technology: it examines the vibration signature of a rotating machine and infers the approaching failure from the unwanted change. Vibration itself is an undesirable condition, caused by pulsating forces.
Many failure modes are easy to detect by analyzing the spectrum or the waveform, already in the early stage of development (Figure 1): bearing fault, imbalance, misalignment, gearbox and electrical faults. The technique also works the other way round: the vibration signature after installation is an acceptance criterion and quality control for precision work (alignment, balancing).
The heart of the program: FMEA-driven, targeted data
Section titled “The heart of the program: FMEA-driven, targeted data”The program does not begin at the instrument but at the list of failure modes: the FMEA (or RCM/PMO) says what has to be detected in the first place.
“Vibration analysis” as a task assignment is too general. To identify the failure mode, several things must be got right: data type, frequency range, resolution, the type and amplitude of the alarm, the extent of the fault. A general data set with fixed resolution lines and a standard Fmax is of limited value. If the data targets the failure mode, the fault is detected with the longest P–F interval. This is why it is mandatory to involve the vibration analyst (with at least an ISO Level 2 certification) in the FMEA/RCM/PMO process.
Data types
Section titled “Data types”The data type is not decided by taste but by the frequency range of the fault you are looking for.
Figure 2 — the frequency range of the fault decides the data type.
- Velocity: general purpose; imbalance, misalignment, looseness, advanced bearing faults.
- Acceleration: high-frequency faults, for example gearbox faults and rotor bars.
- Displacement: shaft displacement measured with an eddy current probe at low speed.
- Further sensing modes: high-frequency detection (HFD), acceleration enveloping, peak monitoring.
With the wrong data type it is unlikely that the fault will be detected in time.
Spectrum, Fmax and resolution
Section titled “Spectrum, Fmax and resolution”Even a well-chosen data type can be prepared in such a way that the information needed to identify the failure mode is not in it.
Figure 3 — failure modes appear in different parts of the spectrum; the Fmax and the resolution decide whether they are visible at all.
Failure modes sit in characteristic parts of the spectrum: at the sub-harmonics, and in the low, medium and high ranges. It is an established empirical pattern that imbalance shows up at 1× running speed and misalignment at the 2× peak (often with an axial 1× alongside), while bearing and gear mesh frequencies sit in their own bands. These are typical patterns, not rules. If the Fmax (the breakpoint of the spectrum) is too low, the necessary data is left out, because what is not collected cannot be analyzed; if it is too high, the resolution drops. Too little resolution masks the individual peaks, too much lengthens the measurement time.
Alarms and dynamic data collection
Section titled “Alarms and dynamic data collection”This is where two things that often get blurred part company: monitoring the integrity of the asset and managing the failure. As long as it is monitoring, the rhythm is set by the P–F interval; the moment there is a fault, it switches into failure management.
Figure 4 — warning and critical levels, and the tightening of the collection as the fault develops.
- Do not rely on “trailing edge” analysis: manual review limits how many machines one analyst can handle. With the principle of the focused alarm only the exceptions have to be looked at; if the data has not changed, there is nothing to review.
- The level of the overall alarm is set high because of false alarms, so it is too high for the small differences of an incipient bearing fault. Its correct application is tripping the machine, to limit the damage.
- Amplitudes must be backed by data: the ISO/API alarms are guidance only, because the transmission path follows from the design of the machine. First collect data, then set a warning (suspicion) and a critical (confirmed) level.
- The rate of change matters just as much as the momentary value: the trend signals earlier than any absolute threshold.
- Target-value alarm: set the threshold knowing the consequence. If the imbalance of a fan caused by fouling reaches a given mm/s value, the bearings will be destroyed; the alarm is therefore the level of fouling that is still tolerable.
The three dimensions of the alarm structure:
| Dimension | Examples |
|---|---|
| Name of the alarm (target data) | overall; 1–6× harmonics; sub-synchronous; low, medium, high range; gearbox and piston pass |
| Type of the target data and the range | single value, peak, all peaks, synchronous and non-synchronous peaks, band energy; the range is absolute, frequency, magnitude, phase angle or time |
| Activation and side | absolute; % change since the last one, versus the baseline or the average; % above the noise floor; rate of change; high, low or both sides |
Frequency of data collection. By default at half the P–F interval, in the constant phase of the life cycle; with several failure modes it is aligned to the fault with the shortest P–F, and tightened at commissioning. When a fault is detected, the ladder is this: the detected fault is first a “reasonable assumption” (it may fail within two to three months), so data must be collected weekly until it develops; from then on expert help must be called in, because it may fail on any day; if a shutdown cannot be arranged, they may keep running with hourly data collection. The shutdown must be initiated at the transition into rapid development.
On-line and off-line vibration monitoring
Section titled “On-line and off-line vibration monitoring”A machine monitoring system has to work on three functional levels: protection, monitoring, diagnostics. The on-line (installed) system collects continuously from every sensor on the critical machines, handles two alarm levels (alert, danger), trips at the danger level on two radial vibration signals (2oo2), and provides transient data, FFT spectrum, orbit, trend and Bode plot. The off-line program measures cyclically on the less critical machines with a portable analyzer. Tripping the machine therefore belongs to the on-line, protection level.
Industrial and safety context
Section titled “Industrial and safety context”The technique is industry-independent. In a hazardous (Seveso) plant, early fault detection prevents the unplanned shutdown and the secondary damage, while the machine trip triggered by the on-line system limits the consequences of the damage. An important boundary: a vibration trip is machine protection, it does not replace the certified, SIL-rated safety instrumented function.
Data collection has risks of its own: access and escape route, temperature, potential process failure, stored energy, explosion hazard, chemical emission. In the development phase the risk grows.
Putting it into practice
Section titled “Putting it into practice”- Start from the FMEA (or from RCM/PMO): which failure mode needs vibration monitoring.
- Assign a data type, an Fmax and a resolution to the failure mode, and designate the measurement point.
- Take baseline data on a freshly installed machine; this is also the quality control of the installation.
- Set data-backed alarms (warning and critical) with the activation type.
- Handle the frequency dynamically, and switch over to exception-based work.
Hands-on
Section titled “Hands-on”Pick a pump or a fan you know well. Write down its two most likely failure modes, assign a data type, an Fmax and an alarm type to each, then state the collection frequency from the estimated P–F interval.
Common mistakes
Section titled “Common mistakes”The pitfalls are not in the instrument but in the way the program is built:
- General data collection point and general task assignment without a failure mode. Correctly: let the points target the identified failure modes, and let the vibration analyst sit in on the FMEA/RCM/PMO.
- Fixed, time-based interval ignoring the P–F. Correctly: dynamic frequency, by the principle of failure mitigation.
- Standard alarms applied blindly. Correctly: ISO/API is guidance only, the amplitude comes from your own data.
- “Trailing edge” analysis. Correctly: exception-based work, with focused alarms.
- Dropping the context. Correctly: interpret the vibration signal together with the operating parameters, otherwise process-induced cavitation will produce a false prediction.
When NOT to use it (the limits of the method)
Section titled “When NOT to use it (the limits of the method)”Vibration analysis is strong, but not universal.
| Situation | Why not vibration analysis | The right answer |
|---|---|---|
| A problem starting in the lubricating oil (water, abrasive contamination, wear metals) | the metals often show up in the oil first | oil and fluid analysis |
| Stationary equipment, an electrical or thermal fault | there is no rotating excitation | infrared thermography |
| A certified safety function is needed | a vibration trip is machine protection, not an audited protection layer | LOPA and SIL, IEC 61511 |
| No designated failure mode, only “let’s measure everything” | a general data set with fixed resolution | FMEA first |
Take it home (keys)
Section titled “Take it home (keys)”- The list of failure modes is the input, not the instrument. Without an FMEA the program measures blind.
- The rate of change matters just as much as the momentary value. Alarm on the trend too.
- The overall alarm is for tripping the machine, not for early detection.
- The standard is guidance only: the amplitude comes from your own data.
- The frequency is dynamic: half the P–F, then weekly, expert help, finally hourly.
Self-test
Section titled “Self-test”- Why is the overall alarm unsuitable for detecting an incipient bearing fault?
- You are looking for a gear mesh frequency fault in a gearbox: which data type do you choose, and what can you get wrong with the Fmax?
- The program has detected a developing fault. At what interval do you keep collecting, and when do you initiate a shutdown?
Answer key: 1) It is set high because of false alarms, so it masks the small differences. · 2) Acceleration; with too low an Fmax the gear mesh frequency does not even get into the spectrum, with too high an Fmax the resolution is lost. · 3) Weekly until it develops, from there calling in expert help, and if a shutdown cannot be arranged, hourly; the shutdown at the transition into rapid development.
How does this show up in digital practice?
Section titled “How does this show up in digital practice?”The alarm–action–follow-up chain is realized in software too.
| Program element | Digital implementation | What it delivers |
|---|---|---|
| Crossing an alarm | automatic work order (CMMS) | the alarm turns into an action |
| Dynamic collection interval | condition-based (PdM) scheduling | the frequency adjusts to the risk |
| Focused alarm | asset condition dashboard, exception filtering | only the exceptions reach the analyst |
| Operator round observation | digital check round, shift log | the round observation lands next to the vibration data |
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”The observations of the operator round (abnormal noise, heat, leakage), recorded in the shift log (OPEREX), give an early, complementary signal alongside the vibration data. The actions assigned to the alarms, the tightened collection interval and the shutdown decision can be documented in an auditable way.
Terminology (HU / EN / JP)
Section titled “Terminology (HU / EN / JP)”| Hungarian | English (canonical) | Japanese | Note |
|---|---|---|---|
| Rezgésdiagnosztika | Vibration analysis | — | CBM technique |
| Rezgésfigyelés | Vibration monitoring | — | a source concept in its own right |
| Rezgési sebesség | Velocity | — | general-purpose data |
| Gyorsulás | Acceleration | — | high frequency |
| Elmozdulás | Displacement | — | low speed |
| Kiegyensúlyozatlanság | Imbalance / Unbalance | — | typically a 1× peak |
| Egytengelyűség-elállítódás | Misalignment | — | typically a 2× peak |
| Spektrum töréspont | Fmax | — | the upper limit of the range |
| Figyelmeztetés / kritikus | Warning (alert) / Critical (danger) | — | alarm levels |
Not a concept of Japanese origin, so the JP column is empty.
What does vibration analysis detect?
Rotating machine faults in the early stage: bearing fault, imbalance, misalignment, gearbox and electrical faults, looseness.
Why must the program start from the FMEA?
Because “vibration analysis” is too general: the failure mode is identified only by targeted data. The FMEA designates what needs monitoring, so we detect with the longest P–F interval.
What is the Fmax, and why does it matter?
The breakpoint, the upper frequency limit of the collected spectrum. If it is too low, the frequency of the fault is left out; if it is too high, the resolution drops, and it is resolution that separates the peaks.
How often should measurements be taken?
By default at half the P–F interval. When a fault is detected, weekly until it develops, from there calling in expert help, and if a shutdown cannot be arranged, hourly.
Is the overall alarm good for early detection?
Generally not: it is set high, so it masks the small differences of an incipient bearing fault. Its correct application is tripping the machine to limit the damage.
Related concepts
Section titled “Related concepts”asset condition management · oil and fluid analysis · infrared thermography · reliability strategy · bathtub curve · FMEA · preventive maintenance · criticality analysis
Next step
Section titled “Next step”From here it is worth going on:
- FMEA — this is where the list of failure modes is born, the input of the vibration program.
- asset condition management — how vibration, oil and thermography come together into one condition picture.
- reliability strategy — how a failure mode becomes a maintenance tactic.
References / further reading
Section titled “References / further reading”- Reliabilityweb.com — Uptime Elements, Asset Condition Information domain: the vendor-independent frame of the vibration program.
- ISO 20816-1 (the successor of the earlier ISO 10816) — Mechanical vibration: measurement and evaluation of machine vibration.
- ISO 18436-2 — Vibration condition monitoring and diagnostics: training and certification of personnel. The analyst categories, among them Level 2, come from here.
- API 670 — Machinery Protection Systems. The specification of machine protection vibration monitoring.
- ISO 17359 — Condition monitoring and diagnostics of machines: general guidelines.
In practice
The vibration analyst works by the focused-alarm principle and looks only at the exceptions; the shift log records the mechanical observations of the operator round (noise, heat, leakage) and the actions assigned to vibration alarms shift by shift, and documents the tightened data collection after a fault is detected and the shutdown decision in an auditable way.
Learn more: Maintenance →