· SYSTEMS & TECHNOLOGY

AI/ML Systems in GxP

Artificial Intelligence / Machine Learning

AUTOMATION & DATAGMPGVPCSVQMS

AI/ML in GxP is rarely a standalone product; it is a capability embedded inside other regulated systems. The working examples are concrete: automated visual inspection classifying filled units, anomaly detection over chromatographic or process data, case intake and triage in pharmacovigilance, predictive maintenance on qualified equipment, and document intelligence over batch records and regulatory text. What sets the class apart from every other computerised system is where the behaviour comes from: a conventional system does what its code specifies, and testing confirms the specification; a machine-learning model does what its training data taught it, and its logic is a set of learned parameters no requirements document fully describes.

All 18 system classes →

What this page does not claim

A system class is not a product. SPEQ describes what a CTMS or a LIMS is; the vendor directory at /tools lists the products that implement one, and a GAMP category is a property of an implementation, not of a class.

What a AI/ML Systems in GxP actually is

AI/ML in GxP is rarely a standalone product; it is a capability embedded inside other regulated systems. The working examples are concrete: automated visual inspection classifying filled units, anomaly detection over chromatographic or process data, case intake and triage in pharmacovigilance, predictive maintenance on qualified equipment, and document intelligence over batch records and regulatory text. What sets the class apart from every other computerised system is where the behaviour comes from: a conventional system does what its code specifies, and testing confirms the specification; a machine-learning model does what its training data taught it, and its logic is a set of learned parameters no requirements document fully describes.

That shift relocates the regulatory problem. Intended use must be defined tightly enough that performance against it is measurable — not "detect defects" but a stated performance level on a stated defect taxonomy under stated conditions. The training, validation, and test datasets become part of the specification, so their provenance, representativeness, and integrity fall under the same scrutiny MHRA's 2018 guidance applies to any raw data. And the distinction between a locked model, whose parameters are frozen at release, and an adaptive one, which continues to learn in production, determines the whole control strategy: a locked model can be verified like software and changed like software, under change control; an adaptive model makes change continuous and demands controls to match.

The framework landscape has consolidated. GAMP 5 Second Edition (2022) addresses AI/ML directly in its Appendix D11, and ISPE has since published a dedicated GAMP guide on artificial intelligence; FDA's Computer Software Assurance guidance points the assurance effort at risk rather than documentation volume, which suits model verification well — though its stated scope is software used in medical-device production and quality systems, so pharmaceutical deployments borrow the philosophy rather than fall under the guidance. In January 2026 FDA and EMA jointly published guiding principles for good AI practice in drug development — SPEQ's /good-ai-practice page covers them in full. Beyond GxP, ISO/IEC 42001:2023 defines an AI management system standard, ISO/IEC 23894:2023 gives AI-specific risk-management guidance, NIST's AI Risk Management Framework (AI RMF 1.0, January 2023) offers a voluntary govern-map-measure-manage structure, and the EU AI Act (Regulation (EU) 2024/1689, in force since August 2024 with phased application) adds horizontal legal obligations that sit alongside, not inside, GMP.

In practice, deployment discipline matters more than algorithm choice. The credible pattern is risk-proportionate: define the context of use, decide how much of the decision the model owns and where the human stands relative to it, set quantitative acceptance criteria before verification, and register every model version with its training data lineage and measured performance. Production is where the class differs most from conventional software — model performance degrades as the world drifts from the training distribution, so ongoing performance monitoring is not an enhancement but the continued-verification obligation itself, and retraining is a change with impact assessment, verification, and approval like any other.

WHERE THE BOUNDARY ACTUALLY SITS

Not the digital twin. A twin is typically a mechanistic or hybrid simulation of a physical process; ML models may feed it, but the twin's claim rests on process understanding, not learned classification.

Digital Twins owns it →

Not the measurement layer. PAT instruments and their chemometric models produce the in-process signal; this class covers learned models consuming such data to classify, predict, or decide.

PAT owns it →

Not conventional automation. PLC logic, SCADA, and historians execute and record deterministic, specified behaviour; the distinct obligations here begin where behaviour is learned rather than programmed.

Historians, SCADA & PLC owns it →

Not automatically a medical device. When an ML model's output drives clinical decisions it may be software as a medical device under the IMDRF framework and device software lifecycle requirements — a different regulatory pathway with its own obligations.

WHAT IT HOLDS, AND WHAT CROSSES ITS BOUNDARY

CORE RECORDS

  • Intended-use and context-of-use definitions, with quantitative acceptance criteria set before verification
  • Training, validation, and test dataset specifications — provenance, representativeness, and integrity evidence
  • The model version registry: each released version with its data lineage and measured performance
  • Verification evidence per release against the pre-set acceptance criteria
  • Production performance-monitoring and drift-detection records
  • Retraining and redeployment change records, with impact assessment and approval
  • Human-review and override logs where the model operates with oversight

DATA FLOWS OUT

eQMS

Performance degradation and drift alerts raised as quality events, and model changes routed through change control

Safety / PV Database

Auto-classified and triaged adverse-event cases handed to the safety workflow with confidence context for human review

MES / EBR

In-line classifications and predictions — inspection results, anomaly flags — surfaced into execution and review-by-exception

HOW THIS CLASS IS USUALLY VALIDATED

  • SPEQ synthesis: GAMP 5 Second Edition (2022) addresses ML in Appendix D11, and the honest reading is that the classic categories describe the software shell around the model better than the model itself — the category is a property of the implementation, and the learned component needs a data-and-performance lifecycle (intended use, dataset controls, pre-set acceptance criteria, monitored deployment) that no category label substitutes for.
  • The locked/adaptive distinction drives everything: a locked model is verified against held-out test data and changed only under change control; genuinely adaptive deployment demands continuous-control machinery most GxP processes are not yet ready to defend, which is why locked models dominate current practice.
  • Acceptance criteria are set before verification, on data the model has never seen, with the test set's provenance controlled as rigorously as the result — evaluating on training data is the field's equivalent of testing into compliance.
  • Production monitoring is the continued-verification obligation, not an enhancement: input-distribution drift and output-performance metrics are watched, thresholds trigger quality events, and a model running unmonitored is a system running outside its validated state.

SPEQ synthesis, not a rating. This is SPEQ’s reading of how this system class is commonly approached, offered to help you scope your own work. A GAMP category is a property of a specific implementation, not of a product class, and one deployment routinely spans several. It is not a classification service and does not replace your own documented risk assessment.

AI/ML SYSTEMS IN GXP MATURITY — REACTIVE TO ADAPTIVE
  1. Stage 1 · Reactive

    Models arrive as pilots owned by data scientists outside the quality system. Intended use is a slide, training data is wherever it was found, and nothing defines what happens when the model is wrong — so nothing GxP-critical can responsibly run on them.

  2. Stage 2 · Defined

    AI use cases are inventoried and risk-assessed, a policy defines where ML may and may not participate in decisions, and the first regulated deployments are locked models with human review of every output and documented verification.

  3. Stage 3 · Controlled

    A defined model lifecycle operates: dataset provenance controls, pre-set acceptance criteria, versioned releases through change control, and production performance monitoring with thresholds wired to quality events. Model records are inspection-ready.

  4. Stage 4 · Predictive

    Monitoring becomes anticipatory — drift is detected in input distributions before output quality falls, retraining is triggered by evidence rather than calendar, and human oversight is tuned to measured model reliability per context of use.

  5. Stage 5 · Adaptive

    AI governance operates as a management system across the portfolio: every model's context, performance, and data lineage is visible in one register, controls scale with demonstrated risk, and the organisation can defend any model's decision trail to an inspector as readily as a batch record.

SPEQ’s shared five-stage progression, labelled synthesis. It is not the FDA QMM rating scale and not the scored maturity-assessment domains — assess your quality system for those.

WHAT AN INSPECTION PROBES, AND WHERE IT GOES WRONG

INSPECTION SIGNALS

  • Whether a stated intended use with quantitative acceptance criteria exists — and whether production performance is measured against it, not against enthusiasm.
  • Training and test data provenance: where the datasets came from, how they were controlled, and whether the test set was truly independent.
  • How the deployed model version is known and reconstructable — which version made this decision, trained on what data, verified with what result.
  • What happens when the model is wrong: the human oversight design, override records, and whether misclassifications feed the quality system.
  • Whether retraining and redeployment went through change control, or models were updated the way IT updates software.

COMMON RISKS

  • Silent drift: the process, materials, or population changes, the model's world does not, and performance decays with no error thrown.
  • Training data that underrepresents the rare, critical cases — the defect classes or case types the model exists to catch.
  • Evaluation leakage: test data contaminated by training data, producing verification results the production model cannot honour.
  • Automation bias — human reviewers nominally in the loop who in practice ratify whatever the model outputs, hollowing the oversight the risk assessment relied on.
  • Vendor models embedded in purchased systems with no visibility into training data, versioning, or update cadence, leaving the buyer accountable for behaviour it cannot inspect.

WHO WORKS IN IT, AND WHERE IT IS SHAPED

ROLES

  • Data scientist / ML engineer
  • Model risk owner / business process owner
  • CSV / digital assurance specialist
  • Quality assurance for digital systems
  • Human reviewers in the oversight loop
  • Data steward for training datasets

DELIVERY-LIFECYCLE PHASES

01 Concept & feasibility
02 Design & engineering
05 Process validation & PPQ
07 Commercial release & handover
The full delivery lifecycle →

[ POSITION IN THE FRAMEWORK ]

6 OF 7 DIMENSIONS · 25 LINKS

Machine-learning capability embedded in regulated processes — behaviour learned from data, not specified in code, so the training set becomes part of the specification and validation must follow the model, not the software shell.

06 · QUALITY MATURITY — AI/ML SYSTEMS IN GXP, REACTIVE TO ADAPTIVE

L1
Reactive

Models arrive as pilots owned by data scientists outside the quality system. Intended use is a slide, training data is wherever it was found, and nothing defines what happens when the model is wrong — so nothing GxP-critical can responsibly run on them.

L2
Defined

AI use cases are inventoried and risk-assessed, a policy defines where ML may and may not participate in decisions, and the first regulated deployments are locked models with human review of every output and documented verification.

L3
Controlled

A defined model lifecycle operates: dataset provenance controls, pre-set acceptance criteria, versioned releases through change control, and production performance monitoring with thresholds wired to quality events. Model records are inspection-ready.

L4
Predictive

Monitoring becomes anticipatory — drift is detected in input distributions before output quality falls, retraining is triggered by evidence rather than calendar, and human oversight is tuned to measured model reliability per context of use.

L5
Adaptive

AI governance operates as a management system across the portfolio: every model's context, performance, and data lineage is visible in one register, controls scale with demonstrated risk, and the organisation can defend any model's decision trail to an inspector as readily as a batch record.

SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →

07 · REGULATORY & EVIDENCE

GOVERNING STANDARDS · 8

Derived from the 8 standards SPEQ maps to this subject, across 7 regulatory bodies: FDA, EMA, ICH, ISPE, MHRA, IEC, IMDRF.

RECORDS & OBJECTIVE EVIDENCE

  • Intended-use and context-of-use definitions, with quantitative acceptance criteria set before verification
  • Training, validation, and test dataset specifications — provenance, representativeness, and integrity evidence
  • The model version registry: each released version with its data lineage and measured performance
  • Production performance-monitoring and drift-detection records
  • Retraining and redeployment change records, with impact assessment and approval

COMMON INSPECTION FINDINGS

  • No stated intended use with quantitative acceptance criteria measured against production performance
  • Training and test data provenance unknown, or the test set not truly independent of training data
  • The deployed model version not reconstructable — which version decided, trained on what, verified how
  • Retraining and redeployment done the way IT updates software, outside change control
  • Human reviewers nominally in the loop who in practice ratify whatever the model outputs
EVERY CHIP IS A DOOR · WALK THE FRAMEWORK FROM ANY SUBJECTHow SPEQ maps the framework →
PROFESSIONAL · IMPLEMENTATION GUIDE · SPEQ SYNTHESIS

Choosing, validating, and living with AI/ML Systems in GxP

CHECKING ACCESS

Checking your Professional access…

FREQUENTLY ASKED

Can AI be used in a GMP decision?

Yes — no regulator prohibits it, and regulated deployments already run in visual inspection, deviation triage, and process monitoring. What regulators expect is that the decision remains defensible: a defined intended use, verified performance against pre-set criteria, data-integrity controls over the training data, a known and version-controlled model, and human oversight proportionate to the decision's impact. The practical pattern is graduated autonomy — models recommend before they decide, and full automation is earned with monitored performance evidence. An AI-influenced decision an organisation cannot explain or reconstruct is the finding; the technology itself is not.

How do you validate a model that can change?

By deciding first whether it is allowed to. A locked model — parameters frozen at release — is verified against independent test data, deployed under version control, and changed only through change control with reverification: a demanding but recognisable software lifecycle. A genuinely adaptive model that learns in production makes change continuous, so the controls must be too — predefined boundaries on how far behaviour may move, continuous performance monitoring against acceptance criteria, and automatic containment when thresholds are breached. Most current GxP deployments choose locked models with periodic controlled retraining, because that keeps the validated state a meaningful, auditable claim.

What GAMP category is a machine-learning system?

The software shell takes a conventional answer — a purchased inspection platform is configured commercial software, bespoke model code is custom — and GAMP 5 Second Edition (2022) is explicit that the category belongs to the implementation, not the product, with Category 2 retired entirely. But the learned component is the reason Appendix D11 exists: a model's behaviour is set by training data, not specification, so category thinking alone under-scopes it. The assurance that matters is the model lifecycle — dataset controls, independent verification, monitored deployment, controlled retraining — layered on top of whatever category the surrounding software takes. This framing is SPEQ synthesis, not a classification service.

Does the EU AI Act apply to AI used in GxP environments?

It can, but as a parallel obligation rather than a replacement for GxP. The AI Act (Regulation (EU) 2024/1689) is horizontal law, in force since August 2024 with obligations applying in phases, and it regulates AI systems by risk class — AI that is a medical device or its safety component intersects the high-risk provisions through the device framework. A quality-control model inside a manufacturing process answers first to GMP: Annex 11, data integrity, and validation expectations. Prudent organisations map their AI portfolio against both regimes, because conformity with one has never implied conformity with the other.