· INTERSECTION

AI in GMP × Quality Maturity

A machine-learning model does not just need validating — it tests whether the site's validation programme is mature enough to defend a system whose behaviour is learned, monitored, and revised.

All 20 intersections →

What this page does not claim

An intersection covers what happens only where two axes overlap. It does not restate what either parent page says, and it is not a substitute for reading them.

WHAT MEETS HERE

WHAT ONLY EXISTS IN THE OVERLAP

  • A conventional system stays validated by staying the same; a model stays validated by continuing to perform. Only in this overlap does "requalification" stop being a paperwork event and become an ongoing statistical claim the validation programme must be able to make and defend.
  • Change control assumes change is proposed by someone. Model drift is change with no proposer — the world moves and the model's behaviour effectively moves with it. Whether drift becomes a controlled event or an invisible excursion is decided by change-control maturity, not by the algorithm.
  • The controls a model demands — dataset provenance, independent test data, audit-trailed retraining, monitored deployment — are the same controls a mature validation programme already runs, pointed at a new object. A site that cannot evidence them for spreadsheets cannot evidence them for models.
  • Human oversight of a model is an engineered control that automation bias can hollow out silently. Whether a reviewer will actually disagree with a model is a quality-culture property of the organisation, not a feature of the software — and it is measurable.
  • The joint FDA and EMA guiding principles on good AI practice read, at a GMP site, as a description of organisational readiness. Where a site sits on the maturity ladder largely predicts which of those expectations it can already meet.

The validated state becomes a performance claim

Qualification, as most GMP sites practise it, rests on a quiet assumption: a system that passed its testing and has not been changed is still in its validated state. Configuration management and change control are the whole mechanism for keeping that claim true. A machine-learning model breaks the assumption from the other side. Its parameters may be frozen — a locked model changes nothing about itself — but its fitness depends on the resemblance between production inputs and the data it learned from, and that resemblance erodes without any change being made. The validated state of a model is therefore not "unchanged since qualification"; it is "measured performance still within the acceptance criteria set before release".

That inversion is survivable only at a certain validation maturity. A programme whose qualification evidence is executed protocols can demonstrate that a model passed its tests once; it cannot demonstrate that the model still performs today, because it has no machinery for continuous evidence. The nearest established habit is the one FDA's process-validation lifecycle guidance built for manufacturing processes: continued verification that a process remains in a state of control, fed by monitoring, with excursions investigated. Sites that already run that discipline for processes have the organisational muscle a model requires. Sites that treat validation as a project with an end date do not — and the model, not the auditor, will be the first to expose it.

Change control meets a system that changes itself

Change control is built around a proposal: someone describes an intended change, its impact is assessed, and implementation is verified. Drift offers nothing to propose. No document arrives when the incoming material profile shifts, when a camera ages, or when the case mix a triage model sees moves away from its training distribution. The only way drift enters the quality system is through monitoring thresholds wired to quality events — which means the site's deviation and CAPA machinery, not its IT department, is what stands between a degrading model and a silent one. An organisation whose deviation system already struggles with ageing investigations will not handle drift alerts better.

Retraining is the mirror problem: it is a change, and it must go through change control, but the impact assessment looks unlike anything the change system was configured for. The evidence of impact is a data-lineage comparison and a performance comparison on independent test data, not a redlined specification. A mature change programme extends its impact-assessment discipline to cover that shape. An immature one fails in one of two directions — it blocks retraining because the paperwork does not fit, and the model decays in production; or it waves retraining through as a "like-for-like update", and the site loses the ability to say which model version made which decision.

Both failure directions end in the same inspection question: which version of the model was in use on this date, trained on what data, verified with what result? Answering it requires the model registry to be maintained with the same discipline as a batch record — and that discipline is precisely what the validation and qualification maturity domain measures.

Maturity bounds how much autonomy a model may hold

The practical pattern in regulated deployments is graduated autonomy: a model recommends before it decides, and full automation is earned with monitored evidence. What is less often said is that the ceiling on autonomy is set by the organisation, not the model. Under ICH Q9(R1), the rigour of a risk assessment should be proportionate to what rides on it — but the quality of that risk assessment depends on the organisation's risk-management maturity. A shallow assessment, performed by a site whose QRM practice is a template-filling exercise, will grant a model more autonomy than the evidence supports, and will not recognise automation bias, evaluation leakage, or under-represented critical cases as hazards at all.

The same dependency runs through data integrity. Training, validation, and test datasets are part of the model's specification, so the expectations MHRA's 2018 guidance applies to any raw data — attributability, controlled access, audit-trailed processing, protection from undisclosed alteration — apply to them. A site that has genuinely implemented data-integrity governance can point its existing controls at a dataset and defend its provenance. A site that passed its last data-integrity audit on effort rather than system will discover that a model multiplies every weakness: the dataset was assembled from wherever data could be found, nobody can say who curated it, and the test set quietly overlaps the training set. The model did not create these gaps; it made them consequential.

The oversight loop is a quality-culture control

Most GMP deployments keep a human in the loop, and risk assessments lean on that human heavily: the model classifies, a qualified person confirms. The control is real only if the person actually exercises judgement, and automation bias erodes exactly that — reviewers ratify what the model outputs, at a rate that increases with the model's apparent reliability. The loop then still appears in every procedure and every record while providing no independent assurance at all. No software setting fixes this, because it is not a software property. It is the same organisational property that decides whether an operator stops a line: whether disagreement is safe, expected, and acted on.

It is also measurable, which is what makes it a maturity question rather than a sentiment. Override rates, disagreement rates, and what happened after each disagreement are records the system can keep. A reviewer who has never once disagreed with the model is either overseeing a perfect model or not overseeing it — and only one of those is plausible. Mature organisations treat a near-zero override rate as a signal to investigate the control, feed confirmed misclassifications into the quality system as events, and tune the oversight design to the model's measured reliability in each context of use. Immature ones discover the hollowness of the loop from an inspector's question.

Where the joint FDA–EMA principles land at a GMP site

In January 2026, FDA and EMA jointly published guiding principles for good AI practice in drug development; SPEQ covers them in full on the dedicated Good AI Practice page rather than restating them here. What belongs to this intersection is their shape: read at a GMP site, they are less a checklist for models than a description of an organisation — one with defined intended uses, governed data, lifecycle thinking, monitored deployment, and humans positioned to intervene. Those are maturity properties. A site can locate itself against them honestly and will usually find its position matches its validation maturity level almost exactly.

One scope note belongs here because it is so often misapplied. FDA's Computer Software Assurance guidance is frequently invoked to justify leaner, risk-pointed assurance for AI — but its stated scope is software used in medical-device production and quality systems. Pharmaceutical GMP deployments may borrow its philosophy, and the philosophy suits model verification well; they should not cite it as their governing framework. GAMP 5 Second Edition, which addresses machine learning directly, and the site's own predicate-rule obligations under Part 11 and Annex 11 remain the frame the deployment answers to.

FREQUENTLY ASKED

Can a site with a low validation maturity deploy an ML model in GMP?

It can pilot one; it cannot responsibly let one participate in a GMP decision. The controls a defensible deployment needs — dataset provenance, pre-set acceptance criteria, a version registry, monitored performance wired to quality events, retraining under change control — are all functions of the surrounding quality system, and a site that cannot evidence them for its conventional systems cannot conjure them for a model. The honest sequence is to raise validation and data-governance maturity first, deploy a locked model with full human review second, and expand autonomy only as monitored evidence accumulates.

Does retraining a model require revalidation?

It requires controlled change with verification, which is revalidation in substance if not in ceremony. Retraining alters the parameters that determine the model's behaviour, so it passes through change control: an impact assessment comparing training-data lineage, verification against independent test data using the same pre-set acceptance criteria, and approval before deployment, with the new version recorded in the model registry. What it does not require is repeating the qualification of the unchanged software shell around the model. The recurring failure is treating retraining as a routine update because "the system is validated" — the platform being sound says nothing about what the new parameters will do.

How does ICH Q9(R1) apply to a model deployed in manufacturing?

As the instrument that sets the model's permitted role. Risk management decides how much of the decision the model owns, where the human stands relative to it, and how much verification rigour the deployment warrants — with formality proportionate to what the decision affects. The intersection point is that the assessment is only as good as the organisation's risk-management practice: a mature QRM function will identify hazards specific to learned systems, such as drift, unrepresentative training data, and automation bias, and will revisit the assessment as monitoring evidence accumulates. A template-driven assessment grants autonomy the evidence does not support.

Where do the FDA–EMA good AI practice principles fit into a GMP programme?

As a readiness mirror rather than a new rulebook. The joint principles, published in January 2026, describe the properties a trustworthy AI deployment and its host organisation should have; SPEQ's Good AI Practice page covers them in full. At a GMP site they do not displace the binding frame — the predicate rules, Part 11 and Annex 11 for electronic records and signatures, and validation expectations operationalised through GAMP 5 — but they give a site a vocabulary for the organisational gaps a model will expose. Reading them against the site's own maturity assessment is a cheap and unusually honest gap analysis.