Skip to content

Operational risk assessment

≈ 17 min read · 3,302 words

Risk assessment identifies, analyses and evaluates hazards, then eliminates or reduces the risk with controls: risk = probability × consequence.

It consists of four steps that build on one another — hazard identification, risk analysis, risk evaluation and risk control — with continuous monitoring feedback. Its purpose is a safer, healthier and more reliable plant, one that shifts from reactive, task-based working to proactive, risk-based working, and so prevents unplanned events.

kockazat-ciklus-en.svg Figure 1 — the four steps of risk assessment and the five guiding questions. Risk changes, so after the control the monitoring and re-evaluation are continuous.

Risk assessment touches every role accountable for the safety and reliability of a process plant: plant and shift supervisor · process engineer · operator and maintenance technician · HSE and process safety (PSM) specialist · reliability engineer · asset team leader. The daily risk routine and the pre-work LMRA are directly in the hands of the people on shift; the choice of method and the common risk register are the responsibility of management and the specialists.

After reading this article you will be able to:

  • recall the four steps of risk assessment and the five guiding questions
  • calculate risk from the probability × consequence relation, and read the risk level off a risk matrix
  • name the five levels of the hierarchy of control, from the most effective to the weakest
  • decide when to use which method (HAZOP, What-if, LMRA, JSA)
  • explain the ALARP principle as the criterion of tolerable risk
  • Four steps: hazard identification → risk analysis → risk evaluation → risk control, then back to monitoring (risk is not static).
  • The goal is to answer five questions: what can happen and under what circumstances, with what consequences, who may be at risk, how likely it is, and whether the existing control is enough.
  • Reliability = defence: to operate reliably means to defend against financial loss and against accidents (personal injury, asset damage).
  • The big shift: moving from reactive, task-based thinking to proactive, risk-based thinking, which is the axis of the whole reliability program.
  • Choosing the method: HAZOP for high-risk, continuous processes; What-if/HAZID for middle or low risk; LMRA just prior to the work; JSA for the step-by-step analysis of non-routine, hazardous work.
  • Every change goes through MOC and a risk assessment — operation, technology, maintenance.
  • One common register: incidents + operational risks + actions in a single system — “no forgotten actions”, and risk-based resource allocation.

The purpose of risk assessment is the identification, analysis and evaluation of hazards, then the elimination of the hazard or the reduction of the risk level with control methods. This creates a safer, healthier and more reliable workplace. Safety and reliability spring from the same root: to operate reliably means to defend against financial loss, and to avoid accidents involving personal injury or asset damage.

Risk assessment is the key that makes possible the shift from reactive, task-based thinking to proactive, risk-based thinking. This helps to avoid unplanned events, and improves the reliability and availability of the units.

What is the risk assessment process in a process plant?

Section titled “What is the risk assessment process in a process plant?”

Risk assessment is four repeating steps that build on one another: hazard identification, risk analysis, risk evaluation and risk control, then back to monitoring. The process leads from uncovering the hazards, through quantifying and judging the risk, to the intervention, and it creates awareness of the hazards and the risk.

  1. Hazard identificationfinding, listing and characterizing the hazards.
  2. Risk analysis — understanding the base of the hazards and defining the level of risk. Risk is the product of probability and consequence (risk = probability × consequence); the input can be current and historical data, theoretical analysis, published knowledge, experience and the concerns of stakeholders.
  3. Risk evaluationcomparing the estimated risk against given risk criteria in order to determine the significance of the risk. The canonical criterion of tolerability is ALARP (As Low As Reasonably Practicable): the risk must be reduced until the cost of further reduction would be grossly disproportionate to the safety gained by it.
  4. Risk control — the actions implementing the decisions of the evaluation, then monitoring, re-evaluation and compliance with the decisions. Risk mitigation measures must be sought in the order of the hierarchy of control (from the most effective to the weakest).

The hierarchy of control:

  1. Elimination — removing the hazard entirely (e.g. taking the hazardous substance out of the process).
  2. Substitution — replacing the hazardous substance or operation with a less hazardous one.
  3. Engineering control — a physical barrier between the hazard and the person (interlocks, extraction, closed system).
  4. Administrative control — procedure, work permit, training, limited exposure time.
  5. Personal protective equipment (PPE) — the last line of defence, where the residual risk is held only by the equipment worn on the worker.

The logic of the ranking: the higher up you intervene, the less the protection depends on human behaviour, and therefore the more reliable it is. The process is a good one if along the way it answers the five questions shown in the opening figure as well (what can happen, with what consequence, who is affected, with what probability, is the control enough).

How do you measure the level of risk? — the risk matrix

Section titled “How do you measure the level of risk? — the risk matrix”

The level of qualitative risk is given by a risk matrix: the intersection of the probability (from rare to frequent) and the consequence (from very low to very high) marks out a risk value and risk level, and to this a prescribed risk response is assigned (acceptable → tolerable only with measures → intolerable).

kockazati-matrix-en.svg Figure 2 — illustrative 5×5 risk matrix: the intersection of probability and consequence gives the risk level and the risk response that belongs to it. The scale and the thresholds are set by the organisation; the values shown here are illustrative.

A higher risk level carries a stricter response and a higher management level: the tolerable band can typically be held only if the ALARP principle is satisfied, or if the reason for accepting the risk is documented; intolerable risk must be dealt with as a priority, at the highest levels, and in the most severe case the activity must be stopped.

PHA (Process Hazard Analysis) contains several methods; the choice depends on the level of risk and on the situation.

pha-modszerek-en.svg Figure 3 — choosing the method: process-level PHA (HAZOP, What-if/HAZID) and pre-work assessment (LMRA, JSA).

  • HAZOP — for high-risk, continuous processes (the basic method of the process industry; node-by-node, guideword-driven, systematic deviation analysis).
  • What-if / HAZID — for middle or low risk processes (faster, less formal).
  • LMRA (Last Minute Risk Assessment) — a basic tool to check and control hazards just prior to the work execution. The moment of “stop and look around”: has anything changed since the permit was issued?
  • JSA (Job Safety Analysis) — if the device cannot be prepared safely, the safest method of the work is defined in a written operational instruction, and the safety precautions are defined during the JSA. The JSA identifies and controls the hazards associated with each step of a hazardous or non-routine job — a useful addition to the existing permit to work system.

Before maintenance, every device must be prepared safely and risk assessed. Before opening equipment that may contain hazardous substances or hot, pressurised energy carriers, the safest working method must be chosen (shutdown mode, the sequence of loosening the screws, and so on).

One of the tasks of the regular meetings is to focus on risks:

  • the daily meeting is dedicated to the risks coming from maintenance work;
  • the weekly meeting is dedicated to what has changed and what is new in terms of risk in that unit.

All new changes in operation, technology or maintenance must go through the MOC process, and the risk must be assessed. The regular risk assessment is carried out in every unit with the involvement of experienced and responsible colleagues, and training on the system and the processes must be organised for shift leaders and management.

Tracking events — one common risk register

Section titled “Tracking events — one common risk register”

A key element of the program is that everything belonging to risk and to incidents should be in one single database. In practice this is realised by a common risk management software application. The content of the common register:

  • the list of all incidents (replacing the parallel, scattered systems);
  • the list of all operational risks — a harmonized register for HSE and business/operational risk; the risk level clearly visualized for asset owners; “creating a place for risk”;
  • all related actions from risk and incident management — no forgotten actions; comprehensive action tracking by unit / responsible / risk level;
  • risk-based budget allocation: the transparent allocation of limited resources based on a cost-benefit measure of risk; risk-based action prioritisation and budgeting decisions.

The system identifies, assesses, manages, regularly reviews and documents all risks associated with all activities. The common register that results covers every unit and site: HSE incident management (fires, injuries) and asset failures are in one system. The recommendations of FMEA (Failure Mode and Effects Analysis) are also collected here, to give better support to preventive and predictive maintenance. All of this serves a single goal: the shift from reactive thinking to proactive, risk-based thinking.

Complex risk matters that touch several functions are handled by the program with cross-functional teams (e.g. a corrosion, ESD or LOPC team). The team is made up of the specialists with the best view of the problem, with a clear charter, scope and team goals; in detail in the program hub.

In a Seveso-classified plant, risk assessment is a matter of law and of life and death, not a formality. The process hazards (the scenarios uncovered by HAZOP) are the inputs of LOPA/SIL, for sizing the layers of protection; while the pre-maintenance LMRA/JSA covers the immediate workplace hazard. Together the two close the risk chain: from the designed layers of protection (SIF, PSV, critical alarm) to the momentary workplace control (shutdown mode, purging, isolation). Near-misses and incidents go into the same register, so the lesson feeds back into the risk assessment. Improving reliability may never override risk control: the LMRA or MOC skipped for the sake of “speeding things up” is the typical accident precursor.

How would you run a risk assessment in practice?

Section titled “How would you run a risk assessment in practice?”

A risk assessment rests on a repeating cycle: identify, evaluate, reduce, prioritise, implement, then review. A shift supervisor or an asset team can start it tomorrow along the following roadmap:

  1. Prepare, and appoint the team. Choose the unit or the job to be examined, and call together the specialists with the best view of it (operations, maintenance, process engineering, HSE) with up-to-date process safety information.
  2. Identify and list the hazards. Go through what can go wrong, and record each of them.
  3. Analyse and score the risk. For every hazard estimate the probability and the consequence, and read the level off the risk matrix.
  4. Compare against the criteria, and decide on the intervention. Above the tolerable level look for measures in the order of the hierarchy of control (elimination first, PPE last), keeping the ALARP principle in mind.
  5. Prioritise, and assign actions. Assign to each of them what will be done, by whom and by when, and put it into a tracked register.
  6. Implement, then re-evaluate. After the measure, check whether the risk has fallen, and repeat the cycle regularly.
  • A one-off risk assessment, then onto the shelf — risk changes, monitoring and re-evaluation are mandatory.
  • Skipping the pre-work LMRA “out of routine” — it is precisely the changed circumstance that goes unnoticed.
  • Choosing the wrong method — an expensive HAZOP for low risk, a superficial What-if for high risk.
  • Forgotten actions — if the actions are not in a tracked register, the risk assessment lives on paper but not in reality.
  • A change without MOC — an unassessed modification is the most common hidden source of risk.
  • Siloed risk — if HSE and asset risk sit in separate systems, the asset owner does not see the whole picture.

When NOT to use it? (the limits of the method)

Section titled “When NOT to use it? (the limits of the method)”

Risk assessment is the basic method of process safety, but on its own it is not the right answer to every situation:

  • It does not replace a certified safety function. A qualitative risk matrix tells you that a layer of protection is needed, but demonstrating the safety integrity level (SIL) requires the semi-quantitative or quantitative analysis of LOPA/SIL; the matrix on its own does not size a safety instrumented function.
  • The wrong method is worse than none. A superficial What-if for a high-risk continuous process, and an oversized HAZOP for low risk, are both wasteful and misleading; the method must be matched to the level of risk.
  • A qualitative matrix can suggest false precision. The estimation of probability and consequence rests on judgement; if you treat it as an exact number, the decision looks better founded than it is. Where the consequence could be severe off-site as well, a quantitative estimate is needed.
  • A one-off, “shelved” assessment gives no protection. Risk changes; without monitoring and re-evaluation the analysis goes stale.
  • It does not work without data and experience. Without up-to-date process safety information and people who know the process, the scoring is done blind.
  • Calculate, don’t just sense it: risk = probability × consequence; the matrix is what makes this comparable and rankable.
  • Start the hierarchy of control from the top: elimination and substitution first, personal protective equipment only at the very end.
  • Match the method to the risk: HAZOP for high-risk processes, What-if for smaller ones, LMRA and JSA just prior to the work.
  • Let no action be lost: put the identified risks and actions into a common, tracked register (what / who / by when).
  • Risk assessment is a living process: every change goes through MOC, and the cycle repeats regularly.
  1. The consequence of a hazard is “high”, its probability “possible”. How do you get the risk level from this?
  2. Of the five levels of the hierarchy of control, which is the most effective, and which is the last, weakest line of defence?
  3. What does the ALARP principle mean, and at which step do you use it?

Answer key: 1) On the risk matrix, the intersection of the “possible” column and the “high” consequence row gives the risk value and risk level, and the prescribed risk response belongs to it. · 2) The most effective is the elimination of the hazard, the last is personal protective equipment (PPE); between them come substitution, then engineering, then administrative control. · 3) ALARP (as low as reasonably practicable) is the criterion of tolerable risk; during risk evaluation (step 3) it decides whether the risk is already low enough.

How does this show up in digital practice?

Section titled “How does this show up in digital practice?”

The principle of risk assessment does not stop at paper or at the workshop: the same logic is recorded automatically and becomes trackable in a digital way of working. The mechanism differs, the principle is the same.

Element Digital implementation Value
Risk register Electronic risk register with a visual risk level In one place, always up to date; the owner sees the whole picture
Pre-work LMRA A digital, mandatory LMRA checklist tied to the work permit The work does not start without the check being completed
Identified risk and action Automatic action tracking with an owner and a deadline Not a single action is lost, an auditable trail is created
Change An approval (MOC) workflow with a built-in risk assessment An unassessed change cannot go through
Incident and near-miss A common database that feeds back into the risk assessment The lesson updates the risk picture at system level

The daily routine of risk management naturally lives in the shift log. The agenda of the daily (maintenance risks) and weekly (changes) risk meeting, the result of the pre-work LMRA, the identified risks and the actions assigned to them (what / who / by when) can be recorded shift by shift and retrieved in OPEREX. This way the log delivers what we expect of a common register: that not a single risk action be lost, that the status be visible at a glance, and that an auditable trail be created for the work permit and for risk-based resource decisions.

Hungarian English (canonical) Note
kockázatértékelés risk assessment the overall process
veszély-azonosítás hazard identification step 1
kockázatelemzés risk analysis step 2
kockázat-kontroll risk control step 4
folyamat-veszélyelemzés PHA (Process Hazard Analysis) the family of methods
munka előtti kockázatértékelés LMRA (Last Minute Risk Assessment) just prior to the work
munkabiztonsági elemzés JSA (Job Safety Analysis) step by step, non-routine work
változáskezelés MOC (Management of Change) for every change
kockázati nyilvántartó risk register common, visualized
What are the four steps of risk assessment?

Hazard identification (finding and characterizing the hazards), risk analysis (understanding the level of risk from data and experience), risk evaluation (comparison against the criteria, determining the significance) and risk control (action, then continuous monitoring and re-evaluation).

What is the difference between HAZOP, What-if, LMRA and JSA?

HAZOP is a systematic process analysis for high-risk, continuous processes; What-if/HAZID is a faster method for middle-low risk. The LMRA is the “last minute” check before the work starts, and the JSA is the step-by-step hazard analysis of non-routine, hazardous work — these last two belong to the work execution, not to the design of the process.

What does the shift from "reactive to proactive, risk-based thinking" mean?

That we do not wait for the failure or the event (reactive), but assess and control the risk in advance so that the unplanned event does not occur at all (proactive). This is the greatest benefit of risk assessment, and the basis of improving reliability.

Why should all risks be put into one common register?

So that not a single action is lost, so that the asset owner sees the whole (HSE + business) risk picture, and so that limited resources can be allocated transparently on a risk basis. HSE incidents and asset failures can then be analysed together, in one system.

production reliability program · HAZOP · LOPA and SIL · management of change · near-miss · ESD systems · alarm management · autonomous maintenance · the standard operator round

  1. Start with process-level hazard identification: HAZOP — how it systematically uncovers process deviations.
  2. From there move on to sizing the layers of protection: LOPA/SIL — how an uncovered scenario becomes a certified safety function.
  3. Finally tie it to daily operation: MOC for the risk of changes, and near-miss for feeding the lesson back.
  • ISO 31000 — Risk management. Guidelines: the international standard of the general framework of risk assessment.
  • IEC 61511 — Functional safety in the process industry: the lifecycle of safety instrumented systems, for sizing risk reduction.
  • Directive 2012/18/EU (Seveso III) — the control of major-accident hazards involving dangerous substances.