Managing human failures - HSE inspector's toolkit
- Publisher
- HSE · UK Health and Safety Executive
- Type
- Toolkit
- Date
- Unknown
- Themes
- Human FactorsProcess Safety
Summary
HSE inspector's toolkit on identifying and managing human failures affecting major accident hazards, describing a qualitative human reliability assessment method.
Summary written automatically from the title and document text.
Themes: human factors, process safety.
Extract from the document (first pages)
Text extracted automatically from the publisher’s PDF so it can be searched. Layout, tables and figures are lost and the extract stops after the first pages; read the document itself at HSE.
CORE TOPICS Core topic 3: Identifying human failures
Introduction
Human failures are often recognised as being a contributor to incidents and accidents, and therefore this section has strong links to the section on accident investigation. Although the contributions to incidents are widely accepted, very few sites will proactively seek out potential human performance problems. Human failure is described fully in chapter ‘Introduction to Human Factors’, where different types of human failures are outlined. In summary, there are two kinds of unintentional failures - physical errors (‘not doing what you meant to do’) and mental errors, where you do the wrong thing believing it to be right (i.e. making the wrong decision). In addition, there are intentional failures or violations – knowingly taking short cuts or not following known procedures.
This will be a relatively new area for many dutyholders and so evidence may not be available to demonstrate that a human factors risk assessment has been completed. Therefore, the inspection will be more likely providing guidance on what is expected in such an assessment on COMAH sites. To assist in this process, a description of a method for identifying and managing human failures is attached below. However, some dutyholders will have partially addressed these issues in an unsystematic manner and the question set will tease out the aspects that they have addressed in part.
Most companies, even if aware of ‘human failure’, will still focus on engineering reliability. It is useful to make this point to dutyholders by asking how they ensure the reliability of an alarm in the control room – usually a detailed and robust demonstration will be made, referring to redundancy, testing etc. However, asking them how they ensure the reliability of the operator who is tasked with responding to the alarm will usually reveal some gaps. You may wish to probe how they know that the operator will always respond in the correct manner, and then discuss what factors may effect an inappropriate response (such as tiredness, distractions, overload, prominence of the alarm indication etc). If any factors are identified, you can ask the site how they could be improved (e.g. providing auditory as well as visual indication, providing a running log of alarms). This process is essentially a human reliability assessment and it is useful to talk through this process so that the company is clear what we mean by addressing human failures.
In assessing human performance, it is all too easy to focus (sometimes exclusively) on the behaviour of front line staff such as production operators or maintenance technicians. The site should be made aware that such focus is undesirable and unproductive. There may be management/organisational failures that have the potential to influence several front line human failures (for example, inadequacies in competency assurance). The technique outlined below can be applied to the identification of failures at the management level.
Human failures in Major Accident Hazards
It should be stressed to the site that we are concerned with how human failures can impact on major accident hazards, rather than personal safety issues.
There are two important aspects to managing human failures in the safety critical industries. First, individual human failures that may contribute to major accidents can be identified and controlled. Second, consideration needs to be given to wider issues than individual human error risk assessments; and this includes addressing the culture of an organisation. Positive characteristics that will support interventions on human failures include open communications, participative involvement of all staff, visible management commitment to safety (backed up by allocated financial, personnel and other resources), an acceptance of
underlying management / organisational failures and an appropriate balance between production and safety. These characteristics will be manifested through a strong safety management system that ensures control of major accident hazards.
Human reliability assessment
The information below is intended to assist in the first of these aspects – an assessment of the human contribution to risk, commonly known as Human Reliability Assessment (HRA). There are two distinct types of HRA:
• qualitative assessments that aim to identify potential human failures and optimise the factors that may influence human performance, and
• quantitative assessments which, in addition, aim to estimate the likelihood of such failures occurring. The results of quantitative HRAs can feed into traditional engineering risk assessment tools and methodologies, such as event and fault tree analysis.
There are difficulties in quantifying human failures (e.g. relating to a lack of data regarding the factors that influence performance); however, there are significant benefits to the qualitative approach and it is this type of HRA that is described below. The company should be informed that our expectation is that they conduct qualitative analyses of human performance – identifying what can go wrong and then putting remedial measures in place.
At the end of the visit, it is expected that the company will be left with a human failure risk assessment proforma, together with guidance on its completion. Agreement from the company should be obtained to undertake such analyses on safety critical operations.
Example of a method to manage human failures
The following structure is well-established and has been applied in numerous industries, including chemical, nuclear and rail. Other methods are available, but these tend to follow a similar structure to that described below. This approach is often referred to as a ‘human- HAZOP’, and this is a useful term to help dutyholders understand our expectations. A proforma for recording the assessment of human failures is provided at Table 1.
Overview of key steps
• Step 1: consider main site hazards;
• Step 2: identify manual activities that affect these hazards;
• Step 3: outline the key steps in these activities;
• Step 4: identify potential human failures in these steps;
• Step 5: identify factors that make these failures more likely;
• Step 6: manage the failures using hierarchy of control;
• Step 7: manage error recovery.
Step 1: consider main site hazards
Consider the main hazards and risks on the site, with reference to the safety report and/or risk assessments.
Step 2: identify manual activities that affect these hazards
Identify activities in these risk areas with a human component. The aim of this step is to identify human interactions with the system which constitute significant sources of risk if human errors occur. For example, there is more opportunity for human performance failures in chlorine bulk transfer than there is in a chlorine storage due to the number of manual operations. Human interactions which will require further analysis are:
• those that have the potential to initiate an event sequence (e.g. inappropriate valve operation causing a loss of containment);
• those required to stop an incident sequence (such as activation of ESD systems) and;
• actions that may escalate an incident (e.g. inadequate maintenance of a fire control system).
Consider tasks such as maintenance, response to upsets/emergencies, as well as normal operations. It is important to note that a task may be a physical action, a check, a decision- making activity, a communications activity or an information-gathering activity. In other words, tasks may be physical or mental activities.
Step 3: outline the key steps in these activities
In order to identify failures, it is helpful to look at the activity in detail. An understanding of the key steps in an activity can be obtained through talking to operators (preferably walking through the operation) and review of procedures, job aids and training materials as well as review of the relevant risk assessment. This analysis of the task steps establishes what the person needs to do to carry out a task correctly. It will include a description of what is done, what information is needed (and where this comes from) and interactions with other people.
Step 4: identify potential human failures in these steps
Identify potential human failures that may occur during these tasks – remembering that human failures may be unintentional or deliberate. Consider the guidewords below for the key steps of the activity. Key steps to consider would be those that could have adverse consequences should they be performed incorrectly.
A task may: Not be completed at all (e.g. non-communication);
Be partially completed (e.g. too little or too short);
Be completed at the wrong time (e.g. too early or too late);
Be inappropriately completed (e.g. too much, too long, on the wrong object, in the wrong direction, too fast/slow);
or, Task steps may be completed in the wrong order;
The wrong task or procedure may be selected and completed;
Additionally, there may be:
A deliberate deviation from a rule or procedure (a ‘procedural violation’).
A more detailed list of ‘error types’, similar to HAZOP guidewords, is provided at the end of this section. Note that an operator may make the same failure on several occasions, known as dependency. For example, an operator may miscalibrate more than one instrument because they have made a miscalculation.
Step 5: identify factors that make these failures more likely
Where human failures are identified above, the next step is to identify the factors that make the failure more or less likely.
Performance Influencing Factors (PIFs) are the characteristics of people, tasks and organisations that influence human performance and therefore the likelihood of human failure. PIFs include time pressure, fatigue, design of controls/displays and the quality of procedures. Evaluating and improving PIFs is the primary approach for maximising human reliability and minimising failures. PIFs will vary on a continuum from the best practicable to worst possible. When all the PIFs relevant to a particular situation are optimal, then error likelihood will be minimised.
Some PIFs that should be considered when assessing an activity/task are outlined in the previous section on accident investigation. HSG48 also lists often-cited causes of human failures in accidents under the three headings of Job, Individual and Organisation. These ‘root causes’ of accidents are in effect the factors that can influence human performance and which should be reviewed in a human factors risk assessment. It is important to consider those factors under the control of management (such as resources, work planning and training) as they can often influence a wide range of activities across the site.
Step 6: manage the failures using hierarchy of control
In order to prevent the risks from human failure in a hazardous system, several aspects need to be considered.
• Can the hazard be removed?
• Can the human contribution be removed, e.g. by a more reliable automated system (bearing in mind the implications of introducing new human failures through maintenance etc)?
• Can the consequences of the human failure be prevented, e.g. by additional barriers in the system?
• Can human performance be assured by mechanical or electrical means? For example, the correct order of valve operation can be assured through physical key interlock systems or the sequential operation of switches on a control panel can be assured through programmable logic controllers. Actions of individuals should not be relied upon to control a major hazard.
• Can the Performance Influencing Factors be made more optimal, (e.g. improve access to equipment, increase lighting, provide more time available for the task, improve supervision, revise procedures or address training needs)?
Step 7: manage error recovery Should it still be possible for failures to occur, improving error recovery and mitigation are the final risk reduction strategies. The objective is to ensure that, should an error occur, it can be identified and recovered from (either by the person who made the error or someone else such as a supervisor) – i.e. making the system more ‘error tolerant’. A recovery process generally
follows three phases: detection of the error, diagnosis of what went wrong and how, and correction of the problem.
Detection of the error may include the use of alarms, displays, direct feedback from the system and true supervisor monitoring/checking. There may be time constraints in recovering from certain errors in high-hazard industries, and it should be borne in mind that a limited time for response (particularly in an upset/emergency) is in itself a factor that increases the likelihood of error.
Specific documents
In addition to the general documents that should be requested prior to the visit (see chapter ‘Aim of the Guidance’) it is recommended that the following documents, which are specific to this topic, should also be requested:
• risk assessment documents outlining the main hazards on site;
• any analyses or documentation referring to safety critical tasks, roles or responsibilities.
A Classification of Human Failures
This list of failures, akin to HAZOP guidewords, can be used in place of the simplified version in Step 4 of the method above.
Action Errors
A1 Operation too long / short
A2 Operation mistimed
A3 Operation in wrong direction
A4 Operation too little / too much
A5 Operation too fast / too slow
A6 Misalign
A7 Right operation on wrong object
A8 Wrong operation on right object
A9 Operation omitted
A10 Operation incomplete
A11 Operation too early / late
Checking Errors
C1 Check omitted
C2 Check incomplete
C3 Right check on wrong object
C4 Wrong check on right object
C5 Check too early / late
Information Retrieval Errors
R1 Information not obtained
R2 Wrong information obtained
R3 Information retrieval incomplete
R4 Information incorrectly interpreted
Information Communication Errors
I1 Information not communicated
I2 Wrong information communicated
I3 Information communication incomplete
I4 Information communication unclear
Selection Errors
S1 Selection omitted
S2 Wrong selection made
Planning Errors
P1 Plan omitted
P2 Plan incorrect
Violations
V1 Deliberate actions
Table 1: Proforma for recording identification of human failures Practical suggestions as to how to prevent the error from occurring are detailed in this column, which may include changes to rules and procedures, training, plant identification or engineering modifications.
Not all human errors or failures will lead to undesirable consequences: There may be opportunities for recovery before reaching the consequences detailed in the following column. It is important to take recovery from errors into account in the assessment, otherwise the human contribution to risk will be overestimated. A recovery process generally follows three phases: detection of the error, diagnosis of what went wrong and how, and correction of the problem.
Human Factors Analysis of Current Situation Human factors additional measures to deal NOTES with human factor issues Task or task step Likely human failures Potential to recover Potential consequences Measures to prevent the Measures to reduce the Comments, description from the failure before if the failure is not failure from occurring consequences or improve references, consequences occur recovered recovery potential questions Task step 1.2 – CRO Action Too Late: CR supervisor initiates Emergency shutdown not Optimise CR interface so that Recovery potential would be initiates emergency Task step performed too emergency response initiated, plant in highly operator is alerted rapidly and improved by ensuring that the response (within 20 late, emergency response unstable state, potential provided with info required to CCR is manned at all times and by minutes of detection) not initiated in time for scenario to escalate make decision; training; clear definition of responsibilities practice emergency response Task step 1.3 – CRO Check Omitted: Supervisor may detect Emergency shutdown not Improve feedback from CR Ensure that training covers the checks that emergency Verification not performed that shutdown not initiated, or only partially interface possibility that shutdown may only response successfully completed complete, as above be partially completed. shut down the plant Ensure that the supervisor performs check Task step 1.4.1 - CRO Wrong information Outside operator Delay in performing Provide standard Correct labelling of plant and informs outside operator communicated: provides feedback to required actions to communication procedures to equipment would assist outside of actions to take if CRO before taking complete the shutdown ensure comprehension operator in recovering CRO’s error partial shutdown occurs CRO sends operator to action Provide shutdown checklist for wrong location CRO
This column records the types This column details suggestions as to This column provides the facility to Task steps taken from of human error that are This column records the how the consequences of an incident insert additional notes or comments procedures, walk through of considered possible for this consequences that may may be reduced or the recovery not included in the previous columns operation and from discussion task. It also includes a brief occur as a result of the potential increased should a failure and may include general remarks, or with operators. description of the specific human failure described in occur. references to other tasks, task steps, error. Note that more than one the previous columns. scenarios or detailed documentation. type of error may arise from Areas where clarification is necessary each identified difference or may also be documented here. issue.
Question set: Identifying human failures Question Site response Inspectors view Improvements needed 1 What does the site understand by the term ‘human failure’? Do they recognise the difference between intentional and unintentional errors? 2 Do they consider that human error is inevitable, or can failures be managed, and how? 3 What are the typical ways to prevent human failure?
4 What are the main hazards on the site? How has the site addressed human failures that may contribute to major accidents? (e.g.1 if a significant risk is reactions in batch processes, how has the site addressed human failure in charging incorrect amount or type of product? e.g.2 if a significant risk is transfer between storage and road/rail tankers, how has the site addressed temporary pipework/hose connection failures?) 5 Is there a formal procedure for conducting human failure analyses? – Is there any science/method to how they assess human failures, or is it seen as ‘common sense’? 6 Does the site identify those manual operations that impact on major accident hazards? (for example, maintenance, start-up, shut down, valve movements, temporary connections). 7 Does the site identify the key steps in these operations? – How (e.g. by talking through the task with operators, walking through the operation, reviewing documentation)? – How do they record this analysis/what formal techniques used (if any)? 8 Does the site identify potential failures that may occur in these key steps (e.g. failure to
Links open the HSE publication page or the free PDF on hse.gov.uk; no login is needed.
Crown copyright, reused under the Open Government Licence v3.0, which permits copying and adapting the information with attribution; this site indexes the first pages and links to HSE's own copies, hosting no publisher download files.
Publisher link checked · working