· GAMP 5 · CSV

AI/ML Validation in GxP

Machine-learning and AI systems increasingly touch GxP decisions — visual inspection, deviation triage, predictive maintenance, batch-record review, pharmacovigilance signal detection. They break the assumptions of traditional computerised-system validation because their behaviour is learned from data, can drift over time, and is often not fully explainable. Validating them means governing the data and the model lifecycle, not just the code, while holding the same standard of documented evidence that the system is fit for its intended use.

What an explainer is not

A topic explainer is SPEQ’s synthesis of what a practice involves, cited to the standards that govern it. It does not reproduce their text, and it does not determine which of them apply to your product or process.

[ POSITION IN THE FRAMEWORK ]

7 DIMENSIONS · 26 LINKS

AI/ML systems break classical CSV: behaviour is learned from data, can drift, and is often unexplainable. Validation shifts to governing the data and model lifecycle — and the static-vs-adaptive choice sets the whole regime.

06 · QUALITY MATURITY — AI/ML VALIDATION IN GXP, REACTIVE TO ADAPTIVE

L1
Reactive

An AI tool is deployed on a GxP decision with no validation, no data provenance, and no monitoring.

L2
Defined

A CSV package is produced as if the model were deterministic; training data and drift are not governed.

L3
Controlled

Intended use and risk are classified; a locked model is validated with documented data provenance and defined acceptance criteria.

L4
Predictive

Data drift and performance are monitored in production; retraining is a bounded, pre-authorised, re-validated change.

L5
Adaptive

The AI system runs a named-gate lifecycle — data, model, deployment, monitoring, retirement — with human oversight sized to risk.

SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →

07 · REGULATORY & EVIDENCE

GOVERNING STANDARDS · 4

Derived from the 4 standards SPEQ maps to this subject, across 4 regulatory bodies: FDA, EMA, ICH, ISPE.

RECORDS & OBJECTIVE EVIDENCE

  • An intended-use definition and risk classification for the AI/ML system
  • Training, validation, and test data provenance and representativeness records
  • Performance demonstrated on unseen data against pre-defined acceptance criteria
  • A static-vs-adaptive decision, with a predetermined change protocol where adaptive
  • Ongoing data-drift and performance-degradation monitoring records

COMMON INSPECTION FINDINGS

  • A learned model on a GMP-critical decision with no lifecycle validation
  • Training data provenance, representativeness, or labelling undocumented
  • A continuously learning model changing behaviour outside change control
  • Acceptance criteria defined after the evaluation data were in hand
  • No human oversight or override on a high-impact AI output
EVERY CHIP IS A DOOR · WALK THE FRAMEWORK FROM ANY SUBJECTHow SPEQ maps the framework →

Why classical CSV is necessary but not sufficient

GAMP 5 Second Edition already anticipates AI/ML: it treats the critical-thinking, risk-based, intended-use principles as the framework and adds an appendix on AI/ML considerations. The point of departure from classical CSV is that a deterministic system produces the same output for the same input, so specification-and-verification testing is meaningful. A learned model's behaviour is a function of its training data and can change if the data or the model changes, so the object of validation shifts from "does the code meet the spec" to "is the model, its data, and its lifecycle controlled and fit for the intended use."

This does not discard CSV — access control, audit trails, change control, infrastructure qualification and Part 11 / Annex 11 compliance all still apply, and the revised EU GMP Annex 11 is expected to address AI explicitly. It adds data-lifecycle governance and model-lifecycle governance on top, and it forces an explicit intended-use and risk classification that determines how much of that governance is proportionate.

Static versus adaptive: the decision that sets the regime

The single most important classification is whether the model is locked (static) or continuously learning (adaptive). A locked model is trained, frozen, validated and then behaves deterministically in production — this is the far more validatable case and is where regulated deployments should start. An adaptive model that retrains on live data changes its behaviour without a change-control event, which is incompatible with GMP change control unless the retraining itself is a bounded, monitored, pre-authorised process.

SPEQ synthesis: for GMP-critical decisions, lock the model and treat every retraining as a formal change with re-validation, rather than deploying continuous learning against a live process. Where adaptivity is genuinely needed, the control is a predetermined change protocol that specifies in advance what may change, within what performance bounds, and how it is monitored — the regulatory concept the device world calls a predetermined change control plan is the closest analogue.

Data provenance is half the validation

A model is only as trustworthy as the data it learned from. Validation must establish where the training, validation and test data came from, that it is representative of the production population, that it was not contaminated by leakage between training and test sets, and that labels were assigned by a defensible process. Bias, gaps and mislabeling in training data are defects that no amount of downstream testing will catch, so data provenance and quality become GxP records subject to ALCOA+.

The evaluation set matters as much as the training set. Performance has to be demonstrated on data the model never saw, using metrics tied to the intended use (a false-negative in automated inspection is not equivalent to a false-positive), and with the acceptance criteria defined before testing. Ongoing monitoring for data drift and performance degradation is a required control, because a model that was fit at release can become unfit as the incoming data distribution shifts.

Explainability, human oversight and risk

Risk assessment under ICH Q9(R1) drives how much explainability and human oversight a given use requires. A model recommending maintenance scheduling carries different risk than a model releasing or rejecting product. High-impact GxP decisions generally require a human-in-the-loop, meaningful explainability, and the ability to override, whereas low-impact advisory uses can be governed more lightly. The regulator-facing question is always whether an unexplained or wrong output could reach the patient.

The emerging framework layer — the FDA/EMA good AI practice principles, ISO/IEC 42001 for AI management systems, and the revised Annex 11 — is converging on the same expectations: intended-use definition, data governance, lifecycle management, human oversight, transparency and continuous monitoring. None of it replaces validation; it structures it. SPEQ synthesis: build the AI system's validation as a lifecycle with named gates (data, model, deployment, monitoring, retirement), not as a one-time qualification event.

FREQUENTLY ASKED

Can I validate a continuously learning AI model for GMP use?

Only if the learning itself is bounded and controlled. Uncontrolled continuous learning conflicts with GMP change control because behaviour changes without a change event. The practical path is to lock the model, validate it, and treat every retraining as a formal change — or, where adaptivity is essential, operate under a predetermined change protocol that pre-authorises the bounds of change and monitors them.

Does GAMP 5 cover AI/ML?

Yes. GAMP 5 Second Edition retains the risk-based, intended-use framework and adds specific AI/ML considerations, so the governing methodology is already in place. What AI adds is data-lifecycle and model-lifecycle governance on top of the classical controls.

Why is training data part of the validation record?

Because a learned model's behaviour is determined by its data. Bias, non-representativeness, leakage between training and test sets, or bad labels are defects that testing may not surface. Data provenance, representativeness and quality therefore have to be documented and controlled as GxP records under ALCOA+ principles.

Do Part 11 and Annex 11 still apply to AI systems?

Yes. Access control, audit trails, electronic records/signatures, infrastructure qualification and change control all still apply. The revised EU GMP Annex 11 is expected to address AI explicitly, but the existing computerised-system expectations already bind AI deployments today.

PROFESSIONAL · INSPECTION PLAYBOOK · SPEQ SYNTHESIS

The inspection-readiness playbook for this topic

CHECKING ACCESS

Checking your Professional access…