[ OPERATING INTERSECTION ]

AI in regulated decisions

The operating model for using AI without transferring human authority or losing evidence.

What this page does not claim

SPEQ synthesis for education. Confirm applicable law, current guidance, standards editions, contractual duties, and organization-specific controls before making a regulated decision.

The seam below, its failure modes and its decision boundaries are SPEQ’s practitioner framing — not a regulatory requirement, and not an assessment of any organization.

OPERATING QUESTION

Is the AI evaluated, monitored, and constrained for the exact context in which its output is used?

Capabilities in the same decision

Why this is hard

Assurance is granted to a system; obligation attaches to a use. For every other kind of software those two stay close together, because a chromatography data system is bought to do one thing and does it in one place. An AI capability is bought as a capability and spreads as a habit. It is assessed for a low-consequence purpose — searching an internal library, condensing a long document — approved on that basis, and then, because it is genuinely useful and sitting right there, its output starts appearing in the root-cause narrative of an investigation, in the answer to a supplier questionnaire, in the justification attached to a change. The tool has one owner and one assessment. The use has as many owners as there are people who found it convenient, and it multiplies faster than any governance process moves. That is the asymmetry: what determines the obligation is not what the tool is for but what its output becomes, and the moment of becoming happens at somebody’s desk, silently, with no gate in front of it. The second difficulty runs the other way and is worse for being counter-intuitive. The human review that makes the arrangement acceptable is degraded by the tool working well. When outputs are mostly right, a reviewer’s job drifts from evaluating to confirming, and the two leave an identical record — so the control decays without anything in the evidence changing.

How it fails

Each of these happens with every function doing its own job correctly. That is what makes them seam failures rather than performance problems.

One approval, twenty contexts of use

The assessment on file describes the purpose the tool was approved for, and it is still accurate about that purpose. What it does not cover is the ten others that grew around it, none of which passed a gate because none of them looked like a change — nobody installed anything, nobody configured anything, somebody simply used an available tool for a new job. The governance artifact and the operational reality diverge without any event that either side would have recognised as needing a decision.

Human review decays as the model improves

The control is that a qualified person reviews every output before it is used, and at the start that review is real. As accuracy rises, so does the reviewer’s expectation that the answer is fine; attention drops, and confirmation replaces evaluation. Nothing in the record distinguishes the two, because an attentive review and a signature produce the same artifact — so the effectiveness of the control is unobservable from the evidence meant to demonstrate it. It weakens in proportion to how well the tool performs, which is the opposite of how controls usually behave.

Vendor performance figures describe a different problem

A published accuracy or benchmark number is a real measurement taken against the material the vendor tested on. Your material is internal: house abbreviations, legacy formats, scanned documents, a second working language, two decades of terminology drift. Performance on the vendor’s distribution transfers to yours only by assumption, and without a measurement on a representative sample of your own material there is no baseline — which means there is also nothing against which later drift could be detected.

The output loses its derivation on the way into the record

The answer is copied into a document and becomes part of a record. What does not travel with it is which model version produced it, what it was asked, which sources it was given, and when. Months later the conclusion is questioned and cannot be reconstructed — not because the record is incomplete in the ordinary sense, since it is signed and attributable, but because the record captured the text and not the derivation. The reasoning behind a regulated decision is the one part nobody kept.

What good looks like

The register is keyed by use rather than by tool, and each entry states what the output becomes, who is accountable for the decision it feeds, and what follows if it is wrong — because that consequence, not the technology, sets how much evidence the use needs. Acceptance criteria are written before the evaluation runs, and the evaluation runs on a representative sample of the organization’s own material, so there is a baseline that later measurements can move against. Review is designed to be falsifiable: a proportion of outputs is independently re-checked, or seeded cases with known answers travel the same path, so the effectiveness of the human control is observable rather than asserted. Provenance is captured at generation and travels with the output into whatever it becomes. Adding a new use is treated as a change with a gate in front of it rather than as initiative, and there is a defined way to withdraw one that is no longer holding, because a capability with no exit accumulates uses forever. And the cost of these controls sits in the same business case as the benefit, since a use whose assurance costs more than the time it saves is a tradeoff worth making deliberately rather than discovering.

Who decides what

The person accountable for the decision an output feeds owns that use, and the ownership cannot be delegated to IT, to the vendor or to whoever introduced the tool, because none of them can see what the output becomes. Signature authority never moves: whoever signs a record is accountable for its content regardless of how the draft was produced, and a procedure that lets a model’s output pass into a record without a named human behind it has transferred authority nobody intended to transfer. Quality decides which uses require assurance and what evidence closes them, and decides it per use rather than once per system. Security and the data owner decide the boundary — what may be sent, what may be retained, whether anything may be used to train — and that decision constrains the use rather than following from it. Portfolio and finance own whether the benefit survives the cost of the controls, which is a real decision usually made by not making it. The authority to make explicit before anybody needs it is who may approve a new use: in its absence, new uses are approved by whoever tries one first, and that is the mechanism by which a governed AI programme quietly stops describing what is actually happening.

Questions practitioners ask

Do electronic-records requirements apply automatically because AI was involved?

No. That scoping follows the underlying regulation requiring the record in the first place, not the technology used to produce it. If the output becomes part of a record you are already required to create, retain and sign, the obligations that already governed that record continue to govern it. If it is not such a record, involving a model does not create the obligation. The assessment is therefore of the use and what it produces, never of the tool in the abstract.

Can a vendor’s validation package replace our own assessment?

It can reduce the work; it cannot reach the conclusion. Supplier evidence covers how the product was built and tested against the supplier’s own requirements, and says nothing about your data, your context, your acceptance criteria or the consequence if an output is wrong. Treat it as an input that narrows what you have to establish yourself, and be specific in writing about what it did and did not cover.

Does keeping a human in the loop remove the need to evaluate the model?

No, for two reasons. Human review is itself a control, so its effectiveness is a question with an answer, and a control whose effectiveness is asserted rather than measured is the weakest kind there is. And review is least effective exactly where the model is most fluent and most often right, which is also where an error is least likely to look like one.

How is change control different for an AI capability?

For conventional software the trigger is an act: somebody deployed, configured or upgraded something. An AI capability’s behaviour can move without any act of yours — a hosted model is updated by its provider, a retrieval corpus is re-indexed, the input distribution shifts as your own documents change. So the trigger cannot only be a change you initiated. It has to include periodic re-verification against a held baseline, which is why having that baseline in the first place matters so much.

Critical handoffs

  1. The process owner defines context of use and decision consequence.
  2. Technical teams build data, model, tool, security, and monitoring controls.
  3. Quality and accountable humans retain acceptance, override, and escalation authority.

Shared evidence

  • Context of use, data lineage, and risk assessment
  • Evaluation, security, change, and monitoring evidence
  • Human review, override, incident, and retirement records