· CONTINUITY & RECOVERY

Business Continuity & Recovery

Recovery used to be a contingency and is now the control that decides how bad an incident becomes. In regulated manufacturing it carries a second obligation that IT continuity planning does not anticipate: a restored system must be demonstrably back in a validated state, not merely running, and the records created either side of the gap must be reconcilable. Coming back up is the start of the quality question, not the end of the incident.

What an explainer is not

A topic explainer is SPEQ’s synthesis of what a practice involves, cited to the standards that govern it. It does not reproduce their text, and it does not determine which of them apply to your product or process.

[ POSITION IN THE FRAMEWORK ]

7 DIMENSIONS · 24 LINKS

Continuity for a regulated operation is not only resuming production — it is resuming it with records whose integrity can be demonstrated, which is why an untested restore is an untested compliance position.

06 · QUALITY MATURITY — BUSINESS CONTINUITY & RECOVERY, REACTIVE TO ADAPTIVE

L1
Reactive

Backups run. Whether anything can be restored from them is unknown, and the plan is a document nobody has opened.

L2
Defined

Recovery objectives are stated and a plan exists, but the objectives were chosen by IT without asking what the process and the records actually require.

L3
Controlled

Objectives derive from what the regulated operation needs, restores are tested on a schedule, and the tests include verifying that restored records are complete and attributable.

L4
Predictive

Continuity is exercised against realistic scenarios including a compromised backup, and manual fallback procedures exist and are trained rather than assumed.

L5
Adaptive

Recovery is routine enough to be uneventful, and the organisation can state what it would lose and prove what it would keep.

SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →

07 · REGULATORY & EVIDENCE

GOVERNING STANDARDS · 5

Derived from the 5 standards SPEQ maps to this subject, across 4 regulatory bodies: FDA, EMA, ISO, IEC.

RECORDS & OBJECTIVE EVIDENCE

  • Recovery time and recovery point objectives, with the operational reasoning behind them
  • Restore test records, including verification that restored records are complete
  • Backup isolation arrangements protecting against compromise of the backups themselves
  • Manual fallback procedures for regulated operations, and training against them
  • Exercise records with findings and their resolution

COMMON INSPECTION FINDINGS

  • Backups verified as completing but never restored from
  • Recovery objectives set by IT with no reference to what the process requires
  • Backups reachable from the same credentials as the systems they protect
  • No manual fallback, so a system outage stops regulated operations entirely
  • A continuity plan that has never been exercised against a realistic scenario
EVERY CHIP IS A DOOR · WALK THE FRAMEWORK FROM ANY SUBJECTHow SPEQ maps the framework →

Two frameworks, one requirement

ISO 22301 specifies a business continuity management system: business impact analysis establishing prioritised activities and their recovery time objectives, risk assessment, continuity strategies, documented plans, and — the clause implementations most often skip — an exercise programme that actually tests them. The 2019 edition replaced the 2012 version, and Amendment 1:2024 added climate-action considerations.

EU GMP Annex 11 gets to a narrower version of the same requirement from the regulated side: systems supporting critical processes need documented, tested continuity provisions such as a manual or alternative arrangement, with the time to bring them into use based on risk, and backups whose integrity and restorability are checked during validation and monitored periodically. The two are complementary — Annex 11 makes it mandatory for a class of systems, ISO 22301 provides the enterprise structure — and running them as separate programmes produces two partial answers.

Recovery objectives expressed in hours mislead

A recovery time objective of four hours means little on a manufacturing site. The operationally meaningful statements are how many batches are in flight, what happens to the ones mid-process when the system carrying their records stops, and how long the manual fallback can carry production before the record backlog itself becomes the constraint. A four-hour RTO on a system whose loss stops a sterile fill mid-cycle is a different proposition from four hours on a reporting server.

This is what the business impact analysis is for, and it is where most implementations are weakest — objectives set by IT, in IT units, without the operational question ever being asked. The BIA is worth doing properly precisely because it is the input everything else inherits.

Backups that survive the incident

Ransomware changed the backup requirement, because modern intrusions target backups first and dwell long enough to be present in several generations of them. A backup reachable with the credentials that were compromised is not a backup. The properties that matter are immutability or offline separation, credentials distinct from production administration, and retention long enough to reach back before the intrusion started — which requires knowing when it started.

Restore testing is the other half and is routinely satisfied on paper. Restoring one system on a quiet afternoon proves very little about restoring forty in a defined order, with the network segmented, under an incident, by people who have not slept. The gap between the test that was performed and the recovery that will be needed is where continuity plans fail.

A restore is a change to a validated system

This is the obligation general continuity planning does not carry. A recovered GxP system has to be verified back into its validated state — configuration, interfaces, user access, audit-trail continuity — and the gap between the last known-good backup and the incident has to be reconstructed and reconciled from paper, upstream systems, or acknowledged as lost.

Manufacturing on a paper fallback during the outage is legitimate where that fallback is a defined, trained, previously tested part of the continuity arrangement. It is not legitimate when improvised during the incident, because the records were then produced under an unqualified process. The authorised person or quality unit making the release decision afterwards needs the integrity assessment, not just confirmation that the systems are back.

SPEQ interpretation — exercise the assessment, not only the restore

Most regulated organisations test disaster recovery, and the test measures how long a restore takes. Almost none rehearse the record-integrity assessment, which is the step that actually determines whether product can be released. The predictable result is that recovery goes roughly to plan and then several days are lost while the quality unit invents an assessment method under pressure to release.

Making that method an artefact rather than an improvisation is cheap: a documented procedure naming the GxP systems in scope, how the incident window is established, what corroboration is acceptable per system, and who signs the conclusion. Written in advance it takes a day; written during an incident it takes a week and is written by exhausted people.

FREQUENTLY ASKED

How should recovery objectives be set for a manufacturing site?

From the operation, not from IT convention. Hours alone mislead: the meaningful questions are how many batches are in flight, what happens to those mid-process, and how long the manual fallback can carry production before the record backlog becomes the constraint. That is what the business impact analysis is for, and it is the input every later decision inherits.

What makes a backup survivable against ransomware?

Immutability or genuine offline separation, credentials distinct from production administration, and retention reaching back before the intrusion began — which requires establishing when it began. A backup reachable with the credentials that were compromised is not a backup, and modern intrusions dwell long enough to be present in several generations.

Is restoring a GxP system from backup a validated activity?

The ability to restore is a validation deliverable — Annex 11 requires backup integrity and restorability to be checked during validation and monitored periodically. A production restore after an incident is a change to a validated system: it needs verification back into the validated state and reconciliation of the gap since the last good backup.

Can product made on paper records during an outage be released?

It depends whether the paper process was a defined, trained and previously tested part of the continuity arrangement, or improvised during the incident. Annex 11 expects continuity provisions for systems supporting critical processes to be documented and tested. Where the fallback was qualified in advance the records are usable; where it was invented under pressure, the release decision has to address that the process itself was unapproved.

PROFESSIONAL · INSPECTION PLAYBOOK · SPEQ SYNTHESIS

The inspection-readiness playbook for this topic

CHECKING ACCESS

Checking your Professional access…