Alarm management
≈ 15 min read · 3,020 words
Alarm management is the lifecycle process of designing, operating, monitoring and maintaining the alarm system so the operator faces a manageable alarm load.
Figure 1 — the alarm management lifecycle: design → rollout and operation → monitoring and improvement, with a continuous improvement feedback loop. (It follows the ISA 18.2 / EEMUA 191 lifecycle logic.)
Who is this for?
Section titled “Who is this for?”Alarm management touches every role responsible for the safety and undisturbed running of a process plant: plant and shift manager · control engineer / DCS engineer · operator · process engineer · process safety (PSM) specialist · reliability engineer · asset team leader. The daily alarm picture and the reporting of faulty instruments are in the hands of the people working the shift; the alarm philosophy, the Master Alarm Database and the rationalization are the responsibility of management and the specialists.
Learning objectives
Section titled “Learning objectives”After reading this article you will be able to:
- state what alarm management is, and list the phases of the lifecycle
- explain why human reliability limits the alarm load that can be handled
- recognize the “bad actor” and the stale alarm, and name the most important alarm KPIs
- read the ISA 18.2 / EEMUA 191 benchmark (~1 alarm / 10 minutes, flood threshold)
- explain why an alarm change requires MOC, and how the alarm is tied to the LOPA/SIL layer of protection
What is alarm management?
Section titled “What is alarm management?”Alarm management is the discipline that spans the entire lifecycle of the alarm system: it defines what we call an alarm, how we prioritize, what the operator’s expected response is, and it ensures that the operator responds to a genuinely important, manageable number of alarms. It breaks down into three big phases: (1) design (philosophy, identification, rationalization, detailed design), (2) rollout and operation (implementation, training, operation, maintenance), (3) monitoring and improvement (monitoring and assessment, management of change, audit). The “bad actor” is the point that generates the most unnecessary alarms (often because of a faulty instrument or a badly set limit); dealing with it is the quickest win in reducing the alarm load.
In brief
Section titled “In brief”- Humans are limited: a mass of alarms cannot be handled reliably; the goal is a manageable alarm load.
- An extra layer of protection: a good alarm, together with the operator’s response, can be one of the independent protection layers (IPL) of LOPA, and so contributes to risk reduction.
- Lifecycle: philosophy → identification → rationalization → design → implementation → training → operation → maintenance → monitoring → MOC → audit. Without maintenance it degrades.
- KPIs: “bad actor” (most frequent) alarms, stale (stuck) alarms, most frequent alarms, shelved/suppressed alarms, the number of configured/standby alarms.
- Benchmark: according to EEMUA 191 / ISA 18.2 the guide value for the load in normal operation is on the order of ~1 alarm / 10 minutes / operator; the goal is to avoid the alarm flood.
- MOC is mandatory: an alarm change must also meet the expectations of HAZOP, LOPA and SIL.
- Base documents: the Alarm Management philosophy, the Master Alarm Database, the complete design documentation.
Why it matters
Section titled “Why it matters”The need for alarm management stems from the fundamental limits of human reliability: the highest integrity level cannot be held if the operator has to fight through a mass of alarms. The benefits of better alarm handling: greater process safety and reliability, and also more production and better quality, at lower cost.
Alarm management minimizes the probability of abnormal situations arising, which are caused by: nuisance and flooding alarms, faulty design and logic configuration, malfunctioning devices, duplication, poor MOC, interlock faults or unregulated PID loops. A well-managed alarm system therefore adds a further layer of protection, and so contributes to overall risk reduction.
What are the phases of the alarm management lifecycle?
Section titled “What are the phases of the alarm management lifecycle?”Alarm management is organized around an AM plan, whose phases are: philosophy, identification, rationalization, design, implementation, training, operation, maintenance, monitoring and assessment, management of change (MOC), audit. This is a lifecycle, continuous improvement process: if the alarms and the associated equipment are not regularly maintained, the performance of the system degrades over time.
In practice the lifecycle means the following:
- the regular review of the alarm KPIs: top bad-actor alarms, momentary stale alarms, most frequent alarms, shelved and suppressed alarms;
- assessing the situation, re-evaluating the alarm priorities, reducing the number of configured and standby alarms;
- launching an MOC to redesign faulty alarms and logic — the alarm change must meet the HAZOP, LOPA, SIL expectations;
- reconfiguring the DCS according to the identified database changes;
- a systematic alarm rationalization team on the selected plant;
- establishing dynamic alarm priorities and limits;
- analyzing the alarm picture of abnormal situations (trip, start-up/S-U, shutdown/S-D);
- the routine reporting of faulty devices;
- up-to-date documentation — recording every change in the proper place (Master Alarm Database, P&IDs, process instructions, operating manual);
- a follow-up review that the changes really work well.
Figure 2 — the regularly monitored alarm KPIs and the guide benchmark (EEMUA 191 / ISA 18.2).
Alarm priorities
Section titled “Alarm priorities”Alarm priority is the means by which the relative importance of alarms can be distinguished. The system must present the alarms audibly and visibly by priority, and must make it possible to sort by priority. The two main considerations of prioritization: the potential severity of the consequence (personal, business, environmental), and the time needed between the alarm and a successful intervention. The philosophy defines three audible levels and one non-audible (log/journal) level, with a typical 80/15/5 distribution (low/medium/high):
| Priority level | Description | Target share | Max proportional rate |
|---|---|---|---|
| Highest (emergency / urgent) | immediate operator intervention, moderate-to-severe consequence potential | ~5% | very rare |
| Medium (high / warning) | prompt operator attention; may escalate to a higher priority | ~15% | < 10 / shift |
| Lowest (low / information) | operator attention, information, low-to-moderate consequence | ~80% | < 6 / hour |
| Non-audible (log / journal) | important enough to record, but not an audible event | — | — |
Alarm classes and highly managed alarms (HMA)
Section titled “Alarm classes and highly managed alarms (HMA)”The alarm philosophy sorts alarms into classes, because the class decides what design, documentation and training requirement applies to them:
- Critical from a personal safety viewpoint — the direct protection of human life.
- Fire and gas hazard indication — e.g. gas-detector thresholds (20% LEL, 40% LEL).
- Critical from a process safety viewpoint (tied to the SIS/SIF): the alarm that pre-warns of an interlock, the alarm that indicates the interlock / the “first-out” alarm, the alarm indicating a shutdown (trip), the alarm indicating an unexecuted command of the SIF final element (e.g. a safety shut-off valve), the POS/MOS switch alarm, the discrepancy alarm of the NooM voting unit.
- Environmental — critical from an environmental or corporate regulatory viewpoint.
Safety- and process-safety-critical alarms are typically highly managed alarms (HMA), to which stricter requirements apply: temporary suppression (shelving) with controlled access, a separate Out of Service (OOS) procedure, mandatory initial and refresher training with documentation, mandatory initial and periodic testing, maintenance training, and a mandatory audit. These alarms provide the direct link toward LOPA/SIL and the safety instrumented functions.
How to introduce it
Section titled “How to introduce it”- Leadership commitment and support for the alarm management work.
- A dedicated alarm management team for the given plant.
- Creating the Alarm Management Philosophy document (this is the foundation: what we call an alarm, how we prioritize, what the operator’s expected response is).
- Developing the Master Alarm Database and the complete alarm design documentation.
- Regular performance monitoring: KPIs, faulty alarms and limits (bad actors), reporting of faulty instruments; assessing the situation and re-evaluating the priorities.
- Operator training.
- A periodic audit — it verifies that the design, implementation, rationalization, operation and maintenance of the alarm system are adequate.
At the weekly alarm review the team lists the top 10 most frequent alarms, and sees that a single level switch accounts for a significant part of the whole alarm load: it flips between “high level” and “normal level” every minute (chattering). From the shift log it turns out that the instrument has been indicating faultily for days. The team raises the faulty instrument for repair, and until then, through an MOC, temporarily shelves the alarm with controlled access and an expiry time, then follows up on whether the alarm count really dropped after the repair. Dealing with one single point thus measurably reduces the daily load — and gives the operator back their trust in the system.
Process industry + safety context
Section titled “Process industry + safety context”Alarm handling is a direct instrument of process safety. In the layers-of-protection model (see LOPA/SIL), a critical alarm + the operator’s correct response can be an independent protection layer (IPL) — by the empirical limit it is worth at most an RRF ~10 credit; if the scenario demands a greater reduction, a safety instrumented function (SIF) is needed. This explains why every alarm change has to fit the HAZOP/LOPA/SIL expectations: a “switched off” or suppressed critical alarm punches a hole in the layer of protection. The alarm flood is especially dangerous: precisely in an abnormal situation (trip, shutdown), when the operator would need the clearest possible view, several hundred alarms can flood the screen, and this is prevented by design and rationalization. International practice is set out in the EEMUA 191 and ISA 18.2 standards.
Measurement / audit
Section titled “Measurement / audit”Alarm performance is measurable, and this is the basis of rationalization. The key indicators (Figure 2) are the bad-actor, stale, most frequent and shelved/suppressed alarms, plus the configured/standby count. The quantitative benchmark of ISA 18.2 (which builds on EEMUA 191) gives the acceptable and the still-manageable levels:
| Indicator | Acceptable target | Manageable maximum | Reference |
|---|---|---|---|
| Alarms / 10 minutes (average) | ~1 | < 2 | ISA 18.2 |
| Alarms / hour (average) | 6 | < 12 | ISA 18.2 |
| Alarms / day (average) | 150 | < 300 | ISA 18.2 |
| Alarm flood threshold | — | > 10 alarms / 10 minutes | ISA 18.2 |
| Stale alarms (> 24 h) | 5 | < 10 | ISA 18.2 |
| Peak in the first 10 minutes after an upset | — | < 10 | ISA 18.2 |
| Share of flood periods (10-minute windows) | — | < 1% | ISA 18.2 |
| The 10 most frequent alarms out of the total load | — | < 1% | ISA 18.2 |
The periodic audit closes the loop: it checks whether the whole lifecycle (from design to maintenance) is in order, and whether the system is holding the target values above.
Common mistakes
Section titled “Common mistakes”- A one-off rationalization, then neglect. Why it’s a problem: without maintenance the lifecycle degrades and performance slides back. Instead: regular KPI review and repeated rationalization.
- A system without a philosophy document. Why it’s a problem: there is no shared principle about what an alarm is and what the priority is, so the system becomes inconsistent. Instead: first the philosophy, then the Master Alarm Database.
- “An alarm for everything.” Why it’s a problem: too many configured and standby alarms cause a flood, and the operator ignores them. Instead: only manageable alarms tied to a documented operator response.
- A suppressed/shelved critical alarm without review. Why it’s a problem: a hidden hole in the layer of protection. Instead: HMA rules (controlled shelving, expiry, audit).
- An alarm change without MOC. Why it’s a problem: it bypasses the HAZOP/LOPA/SIL expectations, and carries unassessed risk into the system. Instead: every change through controlled management of change.
- The faulty instrument is not reported. Why it’s a problem: the “bad actor” keeps making noise, and erodes the operator’s trust in the system. Instead: routine, logged reporting of faulty instruments.
When NOT to use it? (the limits of the method)
Section titled “When NOT to use it? (the limits of the method)”Alarm handling is indispensable, but it is not the right answer to every problem:
- An alarm does not replace a certified safety function. With the operator’s response, an alarm is worth at most a risk reduction credit of ~10; if the scenario demands a greater reduction, a SIF (SIL) is needed, not another alarm.
- Do not prioritize “everything as high.” If every alarm is urgent, in reality none of them is; priority works when it holds the 80/15/5 distribution.
- Suppression (shelving) is not a solution, only a temporary bridge. Without controlled access, an expiry and a review, shelving hides the real problem.
- A KPI target is not an end in itself. Reaching the ~1 alarm / 10 minutes benchmark is worth nothing if the remaining alarms are not rationalized and there is no clear operator response behind them.
- It does not work without maintenance. An alarm system left to itself after a one-off introduction returns to the flood state within a few months.
Take it home (keys)
Section titled “Take it home (keys)”- The human is the bottleneck: the goal is not many alarms, but a manageable load the operator can really respond to.
- An alarm can be a layer of protection: a critical alarm + the operator’s response is a risk reduction of ~10; above that a SIF is needed.
- Measure, then rationalize: the bad actor, the stale and the top 10 alarms are the quickest win; the benchmark is ~1 alarm / 10 minutes.
- Prioritize with discipline: an 80/15/5 distribution, with audible and non-audible levels.
- Every change through MOC: an alarm cannot be touched freely, because it may punch a hole in the layer of protection.
- This is a lifecycle: without maintenance and a periodic audit, performance degrades over time.
Self-test
Section titled “Self-test”- Why does a new alarm on its own not solve a critical scenario that requires a large risk reduction?
- What is the guide alarm rate for normal operation according to ISA 18.2, and how many alarms / 10 minutes means a flood state?
- What is a “bad actor” alarm, and why is it the quickest win in reducing the alarm load?
Answer key: 1) Because with the operator’s response an alarm is worth at most a risk reduction of ~10 (RRF) as an IPL; a greater reduction requires a certified safety instrumented function (SIF/SIL). · 2) In normal operation on the order of ~1 alarm / 10 minutes / operator (manageable max < 2); > 10 alarms / 10 minutes is already an alarm flood. · 3) The point that generates the most unnecessary alarms (often a faulty instrument or a badly set limit); since few points give most of the load, dealing with them reduces the alarm count fastest.
How does it show up in digital practice?
Section titled “How does it show up in digital practice?”The principle of alarm handling does not stop at the DCS or the philosophy document: the same logic is recorded automatically and becomes traceable in a digital operation. The mechanism differs, the principle is the same.
| Element | Digital implementation | Value |
|---|---|---|
| Bad actor / stale alarm | Automatic alarm KPI report from the DCS/alarm historian | The most frequent points are visible immediately, no manual collection needed |
| Alarm picture in an abnormal situation | A recorded, replayable timeline of the alarms during a trip/start-up/shutdown | The flood can be analyzed afterwards, the rationalization becomes targeted |
| Reporting a faulty instrument | Log entry with an owner and a status (in the shift log) | The trail of the “bad actor” is not lost, the repair is traceable |
| Alarm change | MOC workflow with a built-in HAZOP/LOPA/SIL check | An unassessed alarm change cannot get through |
| Shelving / OOS | Suppression with controlled access, logged with an expiry and an audit | A switched-off critical alarm does not stay hidden |
Modern digital operating systems realize the same principles as the paper-based alarm philosophy: measuring the alarms, filtering out the bad actors and the controlled tracking of changes — only faster, retrievably and with an audit trail.
Connection to OPEREX (shift log)
Section titled “Connection to OPEREX (shift log)”The operational side of alarm handling lives in the shift log. The operator’s response to the alarms, the alarm picture during abnormal situations (trip, start-up, shutdown), and the routine reporting of faulty instruments are all log entries. The OPEREX shift log thus feeds the performance analysis — which is the recurring “bad actor,” when the flood happened, which instrument fails repeatedly — and leaves an auditable trail of what the operator did about the alarm. This daily feedback fuels the continuous rationalization: the log shows what has to be redesigned next.
Terminology (HU / EN)
Section titled “Terminology (HU / EN)”| Hungarian | English (canonical) | Note |
|---|---|---|
| riasztáskezelés | alarm management (AM) | the whole discipline |
| riasztásfilozófia | alarm philosophy | the base document |
| riasztás-racionalizálás | alarm rationalization | filtering out the surplus |
| „bad actor“ riasztás | bad actor alarm | the most frequent noise-makers |
| beragadt riasztás | stale alarm | persistently active |
| riasztásözön | alarm flood | too many alarms at once |
| mester-riasztás-adatbázis | Master Alarm Database | the unified register |
| polcra tett / elnyomott | shelved / suppressed | temporarily switched off |
| különleges bánásmódú riasztás | Highly Managed Alarm (HMA) | stricter documentation/training/audit |
The terminology follows the wording of the EEMUA 191 and ISA 18.2 standards.
What is alarm management in one sentence?
The lifecycle process of designing, operating, monitoring and maintaining alarm systems, whose goal is that the operator can respond to a genuinely important, manageable number of alarms, so that the alarm system is an extra layer of protection, not a source of noise.
What is a "bad actor" alarm?
The points that go off most often and generate a disproportionate number of alarms (often because of a faulty instrument or a badly set limit). Dealing with the bad actors is the biggest quick win in reducing the alarm load.
How many alarms are "acceptable"?
According to the ISA 18.2 (EEMUA 191) benchmark, in normal operation roughly ~1 alarm / 10 minutes / operator (manageable max < 2), ~6 / hour and ~150 / day; more than 10 alarms / 10 minutes is already an alarm flood. The main goal is to avoid the flood, especially in an abnormal situation (trip, shutdown).
Why does an alarm change require MOC and HAZOP/LOPA/SIL?
Because a critical alarm can be a layer of protection. If, together with the operator’s response, a risk reduction is “credited” to it in the LOPA, then changing or suppressing it weakens the protection, so it may only be modified through controlled management of change, in line with the expectations of the safety analyses.
Related concepts
Section titled “Related concepts”production reliability program · LOPA/SIL · HAZOP · MOC · ESD systems · IOW · operational risk assessment
Next step
Section titled “Next step”- Size the layer of protection: LOPA/SIL — how a critical alarm becomes a certified risk reduction, and where the limit of the SIF lies.
- Tie it to change: MOC — why every alarm change goes through controlled management of change.
- See the wider risk frame: operational risk assessment and IOW — where the alarm fits the operating limits of the process.
References
Section titled “References”- EEMUA 191 — Alarm Systems: A Guide to Design, Management and Procurement. The de facto industry guidance for alarm system design and performance (the target value for the normal operating alarm rate comes from here).
- ISA-18.2 / IEC 62682 — Management of Alarm Systems for the Process Industries. The lifecycle standard of alarm management and the source of the quantitative KPI benchmark.
In practice
The handling of alarms, the alarm picture during abnormal situations (trip, start-up, shutdown) and the routine reporting of faulty instruments can be recorded in the shift log, shift by shift. The OPEREX shift log thus feeds the analysis of alarm performance (which alarm is the »bad actor«, when the alarm flood happened), and leaves an auditable trail of the operator's response — this is the feedback loop of continuous rationalization.
Learn more: Shift log →