Alarm Management & Safety Instrumented Systems
Alarm floods are a documented cause of major industrial incidents: an operator receiving hundreds of alarms during an upset cannot prioritise, and the alarm that mattered is indistinguishable from the ones that did not. Beneath the alarm layer sits a different thing entirely — the safety instrumented function, which acts without the operator. They are independent protection layers, and treating them as one is how the independence a risk assessment assumed quietly disappears.
What an explainer is not
A topic explainer is SPEQ’s synthesis of what a practice involves, cited to the standards that govern it. It does not reproduce their text, and it does not determine which of them apply to your product or process.
[ POSITION IN THE FRAMEWORK ]
7 DIMENSIONS · 21 LINKSAn alarm is a request for a human action within a time limit; a safety instrumented function is a machine action that does not wait for one. Confusing the two produces alarms nobody can answer and protection nobody verified.
06 · QUALITY MATURITY — ALARM MANAGEMENT & SAFETY INSTRUMENTED SYSTEMS, REACTIVE TO ADAPTIVE
Alarms accumulate as they are configured. Operators silence a standing background of them and the meaningful one arrives in the same stream.
An alarm list exists and priorities are assigned, but the priorities were set by whoever configured each alarm rather than by consequence.
Alarms are rationalised: each has a defined consequence, an operator action and a time to respond, and anything failing that test is removed or re-engineered.
Alarm performance is measured — rate, floods, standing and repeating alarms — and the measurements drive changes to the set rather than being reported.
Protective functions are separated from alarms by design, independently verified on a proof-test interval, and alarm load is low enough that every alarm is genuinely actionable.
SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →
07 · REGULATORY & EVIDENCE
GOVERNING STANDARDS · 5
Derived from the 5 standards SPEQ maps to this subject, across 3 regulatory bodies: FDA, EMA, IEC.
RECORDS & OBJECTIVE EVIDENCE
- The alarm rationalisation record: consequence, operator action, and response time per alarm
- Alarm performance measurement over a representative period
- The separation and independence of protective functions from the control system
- Proof-test records for safety instrumented functions, at their defined interval
- Change records for alarm limits and priorities, with the assessment behind them
COMMON INSPECTION FINDINGS
- Alarm rates that make it impossible for an operator to respond to each one
- Standing or permanently suppressed alarms with no record of the decision to suppress
- Alarm limits widened to reduce nuisance without assessing what they were protecting
- Protective functions implemented in the same controller as the control they protect against
- Proof testing of safety functions overdue or never defined
Three layers, and the independence between them
A process is protected by layers: the basic process control system holding it at setpoint, the alarm layer telling an operator something is wrong, and the safety instrumented system acting when the operator will not or cannot. The protection analysis that justified the design credits each of these separately, and that credit is only valid if they fail independently.
Implementing a safety function in the same controller that runs the process removes that independence — a fault or a configuration error can take both out together — which is why IEC 61511-1 expects separation between the safety instrumented system and the basic process control system. The same logic applies to alarms sourced from a device that also performs the trip: one failure, two lost layers.
Rationalisation is the whole of alarm management
IEC 62682 sets out an alarm lifecycle — philosophy, rationalisation, design, implementation, operation, monitoring, change and audit — and rationalisation is the step that is skipped and the one that produces the benefit. The test is narrow: does this alarm have a defined operator response, and is there time to make it before the consequence occurs? Anything that fails is a notification, and configuring it as an alarm degrades every real alarm around it.
Performance is then measured, not assumed. Average alarm rate per operator is the headline number and it conceals the failure mode: the upset condition in which hundreds arrive in ten minutes. Monitoring has to look at peak and flood rates, at standing and stale alarms, and at the alarms that are always present and therefore always ignored.
Safety integrity is derived, never chosen
A safety instrumented function carries a required safety integrity level established by hazard and risk assessment — typically a layer-of-protection analysis — and expressed as a target probability of failure on demand. It is an output of the analysis. An asserted SIL, chosen because it sounded appropriately serious, is unverifiable and usually unachievable by the architecture claiming it.
The lifecycle that follows is specified: a safety requirements specification before design, design and engineering to meet the target with defined proof-test intervals, and operation and maintenance including the proof testing that keeps the claimed integrity real. A safety function whose proof tests have lapsed no longer delivers the integrity the risk assessment credited, which means the risk assessment is no longer valid either.
Bypasses are where this actually fails
Safety instrumented systems rarely fail through design error. They fail because a bypass was authorised verbally during a campaign to get past a nuisance trip, was not recorded, and was still in place months later. Every bypass removes a protection layer for its duration, and an unrecorded one removes it for an unknown duration nobody is tracking.
The control is procedural and unremarkable: authorisation at a defined level, a stated duration, a compensating measure while the bypass is active, and a record of every one with its removal. In a GMP environment there is a second reason to insist on this — a bypass affecting a critical process parameter during manufacture is information about the batch, and it belongs in the batch record rather than only in an engineering log.
SPEQ interpretation — alarm records are GxP records
Alarm systems are managed as engineering assets and their records as engineering records. But an alarm on a critical process parameter, its timestamp, and its acknowledgement are evidence about a batch: they show that an excursion occurred, when, and whether anyone responded. Investigators routinely ask for exactly this and are told the alarm history was overwritten by retention set for operational convenience.
The practical consequence is that alarm configuration change control and validated-state change control are looking at the same object from two directions, and alarm history retention should be set by the record retention the batch requires — not by the historian’s default. Sites that discover this during an investigation cannot retroactively extend a retention period.
FREQUENTLY ASKED
What is the difference between an alarm and a safety instrumented function?
An alarm asks a person to act; a safety instrumented function acts itself without relying on the operator. They are independent protection layers credited separately in a risk assessment, and implementing both in the same controller or from the same device removes the independence that credit assumed.
What makes something an alarm rather than a notification?
A defined operator response and enough time to make it before the consequence occurs. IEC 62682 calls this rationalisation, and it is the step most implementations skip. Anything failing the test is a notification — configuring it as an alarm degrades every genuine alarm around it by adding to the load the operator must triage.
Can a safety integrity level be chosen rather than calculated?
No. A SIL is an output of hazard and risk assessment, typically a layer-of-protection analysis, expressed as a target probability of failure on demand. An asserted SIL with no analysis behind it cannot be verified and is usually not achievable by the architecture claiming it.
Why do bypasses matter so much?
Because every bypass removes a protection layer for its duration, and an unrecorded one removes it for a duration nobody is tracking. Authorisation at a defined level, a stated duration, a compensating measure and a record of every bypass and its removal is the control — and where a bypass affects a critical process parameter during manufacture, it is information about the batch.