Skip to content

Near-miss: record and investigate

≈ 20 min read · 3,916 words

The manhole cover is not in place. A colleague steps towards it, jumps aside, and walks on. Nothing happened, so there is nothing to report — right? Yet the same missing cover, the next day, in a hurry, in poor light, could just as well have been a broken thigh. That is exactly what a near-miss is: the event that only just failed to become harm. Whoever writes it down gets a free lesson about the weaknesses of the system; whoever walks away from it pays the full price later.

A near-miss is an event that ended without injury, release or damage, but under other circumstances would have been a serious accident.

It is at once lagging information about the protective layers that have already weakened, and a leading signal of a future actual event. The purpose of recording and investigating it is not blame, but uncovering the weaknesses of the protective layers before a real accident happens.

near-miss-barrier-en.svg Figure 1 — the near-miss and the accident come from the same situation; the difference is whether the protective layer only just held.

For those who come across near misses and have to decide what to do with them: operator · shift supervisor · plant manager · process engineer · maintenance technician · HSE specialist · reliability engineer.

After reading this article you will be able to:

  • distinguish a near miss from a near-accident;
  • classify an event by severity (PEAR), and recognize the high-potential (HiPo) case;
  • decide what depth of investigation and what team a given event calls for;
  • build a near-miss KPI at Tier 3 level, with a leading dual-assurance partner.
  • A near-miss is an event without consequence but with potentially serious outcome: a free lesson about the weaknesses of the system.
  • A dual-natured indicator: lagging (a barrier weakness that has already happened) and leading (it forecasts) at the same time.
  • Learning only from major events is not enough: those are rare. Near misses are more frequent, so they give a more usable data set.
  • A HiPo near-miss deserves just as rigorous an investigation as an actual accident.
  • The engine of the system is just culture: reporting is blame-free, otherwise near misses never surface.
  • The near-miss is level 3 (Tier 3) of the process safety KPI pyramid.

Because an unreported near miss does not disappear, it only becomes invisible, and stays in the system as a weakened protective layer.

Zero consequence, fatal potential, and a single sentence of cause (the temporary rail was not put back) that can be fixed within days of a report.

What is a near-miss, and where does it come from?

Section titled “What is a near-miss, and where does it come from?”

It is the practical consequence of the realization that a serious accident is almost never the work of a single failure. For a long time the process industry learned from what had already happened. This reactive mode matters, but the strategy of waiting until something happens and then learning from it afterwards is not sufficient on its own: major events are rare and accumulate slowly.

A serious accident is typically the result of several independent protective layers (barriers) being breached at the same time. This is the “Swiss cheese” model, created by professor Jim Reason, and the distinction between active and latent failure belongs to it too: some of the holes have been in the system for a long time, they just had not lined up until now. The near-miss is the moment when one or more barriers have already been breached, but the remaining layers still caught the event.

near-miss-svajci-sajt-en.svg Figure 2 — barrier weaknesses are continuously present; the only question is whether they line up.

The near-miss category includes:

  • a below-threshold release (LOPC, Loss of Primary Containment), for example a spill smaller than a drum;
  • a demand on a safety system: a pressure relief device (PRD) lifting, a safety instrumented system (SIS) or a protective shutdown tripping;
  • an excursion beyond the Safe Operating Limit (SOL), from which a pre-determined action brings the process back;
  • an observation of an unsafe condition, without consequence.

A near-miss ends with “nothing happened”, and that is exactly what makes it hard to recognize. Judging how potentially serious an event was is subjective: it calls for risk estimation, that is, for terrain that is sensitive to cognitive biases.

The two are not the same: the near miss is the broader concept. It does not only mean the possibility of personal injury, but also of occupational illness, damage or loss to people, assets, the environment or the company’s reputation. The near-accident is narrower: we only speak of one if all three conditions are met:

  1. there was a technical failure or an error of action that posed a hazard to the surroundings;
  2. people were present in the hazard zone that arose;
  3. an accident did not happen because the circumstances turned out fortunately.

If these do not hold, the event is not a near-accident: a slight injury without lost working time is already a real accident (NLTI).

near-miss-piramis-en.svg Figure 3 — at the top the four-step severity ladder worked through on the missing manhole cover, below it the ratios of the accident pyramid.

The ladder arranges the same situation by the severity of the outcome: near miss (the missing cover in itself), near-accident (the colleague steps towards it but jumps aside), NLTI (steps into it but only grazes himself), and finally LTI (steps into it and breaks his leg).

The accident pyramid depicts the same logic: under every serious outcome lies a multiple number of milder events. The typical ratios of the DuPont pyramid are 1 fatality : 30 serious injuries : 300 recordable events : 3,000 near-accidents or first-aid cases : 300,000 unsafe acts. Managing the base lowers the apex.

How should a near-miss be recorded and investigated?

Section titled “How should a near-miss be recorded and investigated?”

In a closed loop: low-threshold, blame-free reporting, classification by potential, an investigation whose depth is proportionate to the potential, an action assigned to an owner, and finally the re-measurement of effectiveness and the sharing of the lesson.

near-miss-folyamat-en.svg Figure 4 — the near-miss investigation process, with the two paths that separate at the triage.

1. Make reporting dead simple. The process starts where the event happens: in the shift, at the operator. A card, a mobile form or a shift log entry; many plants run a separate stop card system. Reporting must be blame-free: the purpose of the investigation is never to sanction those responsible.

2. Define plant-specifically what counts as a near-miss (a PRD lifting, an SOL excursion, a below-threshold LOPC, an unsafe condition). Give concrete examples, otherwise everyone will understand something different by it.

3. Classify both the consequence and the potential. Severity must be given in four categories (PEAR: People, Assets, Environment, Reputation), from 0 to 5; the consequence of the highest severity gives the overall classification. The near miss is severity 0, that is, “no consequence” in all four categories.

Severity People Assets Environment Reputation
0 · near miss none none none none
1 slight injury slight damage within the site limited
2 lost working time minor damage single breach local complaint
3 serious, prolonged unit out of operation beyond the site boundary ongoing protest
4 fatality major damage severe damage national
5 multiple fatalities extensive damage massive, persistent international

The actual consequence must be recorded together with the event data, the potential one after the initial investigation. Then comes the key question of the triage: realistically, what would have been the worst credible outcome (realistic worst case) if the remaining protection had also failed? If the answer is a major event, then it is a high-potential (HiPo) near-miss.

4. Run two recording lanes. The HiPo near-miss goes on the same lane as a serious event, typically within one working day into the corporate HSE incident register; for non-HiPo cases the local system is enough. This way the important cases are not lost in the crowd.

5. Match the depth of the investigation and the team to the severity. This is where what the triage decided becomes auditable.

Severity Investigation Team
0–1 simplified analysis no formal team needed
2–3 and every HiPo detailed investigation with structured root cause analysis (Tripod or another methodology) multidisciplinary, preferably business-led, with HSE experts involved
4–5 detailed investigation the team leader and at least one member are independent of the organisational unit concerned

The team must include at least one trained investigator. The mandatory content of the detailed report: the sequence of events, the failed barriers, the root causes and latent failures, the improvement measures with responsible persons and deadlines, and the evidence. Deadline: at most 60 calendar days. As for methodology, the 5 Whys serves quick cause finding, RCA the general case, and Tripod Beta, HFIT and HFACS the barrier-oriented exploration of human and organisational factors.

6. Track the actions for implementation AND effectiveness. An investigation is only worth something if it produces concrete actions with deadlines and owners. Tracking does not stop at closure: it must also be assessed whether the action achieved its goal, and if not, the investigation has to be reopened.

7. Share the lesson. In the form of a learning letter, an awareness note or an action alert; from HiPo and severity 3 upwards with a short summary in English (a “five-pager”: description, consequences, root causes, corrective actions, lessons learnt). Investigation KPIs and trends must be evaluated regularly.

In a process plant (crude oil processing, petrochemicals) the tangible forms of a near-miss are: a PRD or safety valve lifting above the set pressure but below the Tier 1/2 threshold; a SIS or protective shutdown tripping, for example on high-high level; an SOL excursion during start-up, shutdown or normal operation (different SOLs may apply to the different phases); a below-threshold LOPC; and failures of the alarm system.

On the reliability side the “learn from failures” principle is the analogue: frequent, high-quality, prioritized investigation. A typical bad practice is to investigate formally only HSE events, and not the production losses that are near-miss-like in nature. Reliability is driven by the same soft factors as safety: the organisation, performance management and mindset.

Regulatory and standards framework:

  • The Seveso III Directive (2012/18/EU) requires the control of major-accident hazards for establishments handling dangerous substances; part of the safety management system (SMS) is the recording and investigation of events and near misses, and the feedback of the lessons.
  • ISO 45001 explicitly requires the investigation of incidents (including near misses) and corrective actions (clause 10.2), as well as worker participation in reporting.
  • Within the framework of IEC 61511 / IEC 61508 (SIL), activations of safety instrumented systems (demands) mean that a protective layer has been called upon; tracking these feeds back to the demand frequency assumed in LOPA and to the SIL classification.

Putting it into practice: the 20-minute near-miss calibration

Section titled “Putting it into practice: the 20-minute near-miss calibration”

If you want to start tomorrow, begin with this. It needs five or six colleagues, a flipchart and twenty minutes.

  1. Collect (5 minutes). Everyone writes down two situations from the past month in which “something almost happened”. Anonymously, on slips of paper.
  2. Classify (5 minutes). Go through the slips against the PEAR table: what was the actual consequence, and what would have been the worst credible outcome?
  3. Vote on the HiPos (5 minutes). Which slip would have become a major event? Put a red dot on those.
  4. Formulate one action (5 minutes). For the strongest red-dotted case, assign a single action with an owner and a deadline.

Homework: two weeks later, check not only whether the action was closed, but whether it achieved its goal.

The near-miss is level 3 (Tier 3) of the process safety KPI pyramid: more frequent than Tier 1/2 events, and lagging and leading at the same time.

near-miss-tier-en.svg Figure 5 — the process safety event pyramid; the near-miss lives at the Tier 3 level.

Useful indicators:

  • Number of near-miss reports (per month) and its trend: a maturity signal.
  • Ratio of near misses to actual events: high in an alert system.
  • Demands on safety systems: the count of PRD lifts and SIS trips, or the demand rate per system type.
  • Number of SOL excursions (each excursion counted separately).
  • Action closure rate and lead time, and the effectiveness of the closed actions.

In the daily cadence, the near-miss count lives in the Safety column of the performance board: the three usual indicators of the safety block of the SQDC board (Safety, Quality, Delivery, Cost) are LTI, Near Miss and stop card.

Dual assurance. For every critical barrier it is worth pairing one leading (Tier 4) and one lagging (for example near-miss) KPI, and correlating the two; this way it can be tested whether the barrier is getting stronger or weaker. The pattern holds for every critical barrier, not only the alarm system:

Barrier Tier 3 (rather lagging) Tier 4 (leading)
Alarm system number of alarm system failures (from testing, near misses, actual events) percentage against plan of completed alarm tests
Competence of personnel number of near misses, LOPCs and plant trips linked to traineeship, lack of technical understanding or inadequate training percentage of personnel meeting the competence criteria in critical roles
Operating procedure errors due to incorrect or unclear procedures; number of operational shortcuts identified from near misses percentage of procedures reviewed against plan

Most mistakes are not in the method, but in what the organisation does with the report.

Anti-pattern Why it is a problem Good practice
Reporting has a price (even an implicit one) near misses disappear, the number improves, the risk stays a just culture
Reducing the near-miss count is a target figure the metric turns against itself: it encourages under-reporting a rise in reports at the start of the introduction is a good sign
“Nothing happened”, so we close it an underestimated HiPo slips onto the simplified path mandatory realistic worst case estimate, a HiPo threshold
Only HSE events are investigated the barrier weakness hidden in production losses is left out prioritization for every significant loss
No production target, hardly any measurement of lost production there is no reference base a target figure and loss measurement alongside the investigation
“Whoever happens to have time” investigates an untrained investigator, a superficial finding a trained investigator in every team
No structured root cause analysis, the answer is an engineering modification the organisational cause is left untouched RCA methodology, a system cause instead of “the operator made a mistake”
The action is closed, so we are done closure is not the same as a solution effectiveness re-measurement, reopening if the goal was not reached
No dedicated reliability team the barrier weaknesses have no owner named reliability accountability
The lesson stays with the team one team learns, the others do not learning letter, five-pager, sharing of the KPI trend

When NOT to use it (the limits of the method)

Section titled “When NOT to use it (the limits of the method)”
  • It does not replace design risk analysis. HAZOP and LOPA also cover scenarios that have never occurred; the near-miss only teaches about what has already almost happened.
  • On its own it does not forecast the rare, high-impact event. The weak signals of a black swan event often do not arrive in the form of a near-miss.
  • Do not use it for performance appraisal. As soon as reporting carries a personal consequence, the data goes bad.
  • The near-miss rate is not a benchmark. The maturity of reporting culture differs so much between organisations that the count is misleading.
  • Without triage and action, mass recording is administration, not risk reduction.
  • “Nothing happened” is not a closure but a question: what would have been the worst credible outcome?
  • Triage by potential, not by actual consequence: the HiPo near-miss is on the same lane as a serious accident.
  • Culture first, system second. Without blame-free reporting, every step works on distorted data.
  • The barrier is always there in the report: what was breached, and what held.
  • A closed action is not a solved problem. Re-measure the effectiveness, and if it did not achieve its goal, reopen it.
  1. A protective shutdown tripped because of a high level, and there was no release. Is this a near miss, a near-accident, or neither? Justify it with the three conditions.
  2. The actual consequence of a near-miss is zero in all four PEAR categories, while the realistic worse scenario is a fatal fall. What investigation does it call for, and with what team?
  3. In one plant the number of near-miss reports tripled within a year. Is this good news or bad, and with what second indicator would you reinforce your judgement?

How does this show up in digital practice?

Section titled “How does this show up in digital practice?”

The logic of the near-miss does not stop at the paper card: the same principle is realized in software too, in a well-designed incident management system.

Near-miss principle Digital implementation What it delivers
Low-threshold reporting mobile near-miss reporter with few mandatory fields the recording happens where the event does
Potential-based triage mandatory “worst credible outcome” field, automatic HiPo flag “nothing happened” does not close the case
Action tracking and re-measurement action with owner and deadline, then a later effectiveness assessment closure does not cover up an unsolved problem
Sharing the lesson searchable case library, notification to similar plants one team’s lesson becomes the organisation’s

The first and most critical point of the near-miss lifecycle is recording at the source, in the shift where the event happened. This is exactly where knowledge tends to be lost: at the shift handover, in a verbal handover, or because reporting is cumbersome. The OPEREX shift log SaaS records the events of the shift (among them near misses, demands and SOL excursions) in a structured way at the very moment they arise. This way the near-miss is not an entry assembled afterwards from memory, but time-stamped, searchable data: the basis of the investigation and of the Tier 3 KPI trend.

Hungarian English Japanese / note
kvázi esemény, majdnem-baleset near-miss ヒヤリハット (hiyari-hatto): the broader concept
kvázi baleset near-accident the three conditions met together
nagy potenciálú near-miss high-potential (HiPo) near-miss a major event in the realistic worse case
gyökérok-elemzés Root Cause Analysis (RCA) 根本原因分析: structured cause finding
védelmi réteg barrier the layers of the “Swiss cheese”
elsődleges zárás elvesztése Loss of Primary Containment (LOPC) release
biztonsági rendszer igénybevétele demand on safety system a PRD or SIS activation
méltányos kultúra just culture blame-free, honest reporting
súlyossági kategóriák PEAR (People, Assets, Environment, Reputation) a 0–5 scale, near miss = 0
What is the difference between a near-miss and a near-accident?

The near-miss (near miss) is the broader concept: it covers every event that could have caused injury, occupational illness or damage to people, assets, the environment or the company’s reputation. We only speak of a near-accident if it is met at the same time that a technical or action error occurred, that people were present in the hazard zone, and that the accident was avoided only because of fortunate circumstances.

When must a near-miss be classified as a high-potential (HiPo) event?

When the realistically conceivable worst outcome would have been a major event, assuming that the remaining protection had also failed. A HiPo near-miss must be handled on the same lane as a serious accident: rapid recording, detailed investigation with structured root cause analysis, by a multidisciplinary team.

Why is it a problem to treat the number of near-miss reports as a target?

Because rewarding a falling number encourages under-reporting: the figure improves, but the real risk stays, and even becomes invisible. At the start of the introduction a rise in the number of reports is a good sign, because it shows a more mature reporting culture.

How long may a near-miss investigation take?

At most 60 calendar days. For an event of severity 0–1 a simplified analysis is enough; from severity 2 and for every HiPo case a detailed investigation is due, with structured root cause analysis.

How do I know whether anything came of the investigation?

From the fact that the actions are tracked not only for implementation but also for effectiveness. If a closed action did not achieve its goal, the investigation has to be reopened; the bare closure rate is a misleading indicator.

HAZOP · LOPA · SIL · management of change · the 5 Whys · robustness and redundancy · black swan event · safety cross · performance board

If you have understood this, from here it is worth going on — in this order:

  1. the 5 Whys — the engine of the investigation: from the report to an actionable root cause.
  2. LOPA and SIL — how the near-miss becomes a measured demand frequency, and how it feeds back into the sizing of the protective layers.
  3. safety cross — how you make safety performance visible day by day.
  • IOGP/OGP: Process safety — recommended practice on key performance indicators (Report 456). The canonical source of the Tier 1–4 KPI pyramid and of the “dual assurance” barrier pairing.
  • API RP 754: Process Safety Performance Indicators for the Refining and Petrochemical Industries.
  • CCPS (AIChE): Process Safety Metrics — Guide for Selecting Leading and Lagging Indicators.
  • James Reason: Human Error (Cambridge University Press, 1990); A Life in Error (Ashgate, 2013). The “Swiss cheese” model, the distinction between active and latent failure, the concept of just culture.
  • Ronald W. McLeod: Designing for Human Reliability in the Oil, Gas, and Process Industries. Gulf Professional Publishing, 2015.
  • Stichting Tripod Foundation / Energy Institute: the Tripod Beta incident investigation methodology.
  • Gordon, Flin, Mearns: A human factors investigation tool (HFIT) for accident analysis. Safety Science 2005;43:147–171.
  • Shappell, Wiegmann: The Human Factors Analysis and Classification System (HFACS). DOT/FAA/AM-00/7, 2000.
  • Directive 2012/18/EU (Seveso III) and clause 10.2 of the ISO 45001:2018 standard.