· AI GOVERNANCE EVIDENCE FABRIC

Evidence for an AI invocation your quality system can file

8 QUESTIONSSCOPED · LAPSING · VERIFIABLE

A regulated organisation runs an AI tool inside a GxP workflow, and its QA unit is then asked 8 questions about that use — by an inspector, or by its own quality system. SPEQ makes the invocation answer them: what ran, on what, against criteria fixed before the result was seen, with an audit trail. The output is a Governance Evidence Record — scoped to one tool, one version and one declared use, verifiable by recomputation, and it lapses when that system changes.

What this page does not claim

SPEQ structures, cites and documents the evidence for an AI invocation. It does not validate the tool, approve its use, or state that anyone is compliant — validation and release stay the regulated party’s decisions, under their own quality system.

The record is a chain, and one link expires

Each link below is fixed by a hash of its own contents, and the order they were made in is enforced by an append-only chain rather than asserted in a document. That is what makes the second link mean anything at the fifth: criteria frozen before a run cannot be edited after a result is seen, because editing them produces a different criteria id and a different record. The last link is the only one that expires.

THE VERSIONED EVIDENCE CHAIN

  1. Source input

    sha256 per input, with its declared origin

    Every input the run consumed is named, hashed, and has a stated source.

  2. Frozen criteria

    criteriaId — sha256 of the criteria set

    The thresholds were fixed, with written reasons, before the run was submitted.

  3. Execution

    manifestId, plus the pinned components

    What ran, where, against which pinned versions, and how long it took.

  4. Observed metrics

    the response hash the metrics were read from

    What came back. A metric the provider did not return is recorded as missing, never as zero.

  5. Evaluation

    the verdict, bound to that criteriaId

    The observed values judged against the frozen criteria — pass, fail, or indeterminate.

  6. Dependency fingerprint

    jcs-sha256 over the pinned components

    The exact configuration this evidence describes, and which changes to it lapse the record.

  7. Current validity

    the validity block, re-derived on read

    Whether the record still describes the running system. This is the link that expires.

THE 8 QUESTIONS

The bet this layer is built on is that these 8 questions are the same whatever the tool is, and only the thin layer that answers them changes. Each one carries the clauses it answers to; a clause marked binding on a record is one the declared context of use actually brings into scope, and one marked reference is disclosed rather than dropped, because the question still came from it.

Q1

What regulated decision does this output inform?

Intended use is the input to every downstream judgement, and Annex 11 §1 is where the risk management that scales effort to use is stated. Part 11 has no intended-use clause: it governs records once you are in scope, not the question of whether you are.

EU GMP Annex 11 §1

Q2

What is the software category, and what is the criticality?

Categorisation drives validation rigour. Annex 11 §4 and Part 11 11.10(a) are the validation obligations the category scales; §1 is the risk basis for scaling it.

EU GMP Annex 11 §1EU GMP Annex 11 §421 CFR 11.10(a)

Q3

Can this exact run be reproduced?

Reproducibility is what "consistent intended performance" in 11.10(a) means for a stochastic system. Annex 11 §6 accuracy checks are the same obligation applied to computed values.

21 CFR 11.10(a)EU GMP Annex 11 §4EU GMP Annex 11 §6

Q4

Where did every input come from, and is it unaltered?

Annex 11 §5 is the data clause and §7 the storage one; 11.10(b) is the requirement to produce accurate and complete copies, which is unmeetable without provenance.

EU GMP Annex 11 §521 CFR 11.10(b)EU GMP Annex 11 §7

Q5

What were the acceptance criteria, and were they set before the run?

A threshold chosen after seeing the result is not an acceptance criterion. 11.10(f) operational checks and §6 accuracy checks are the clauses a retrofitted threshold defeats.

EU GMP Annex 11 §621 CFR 11.10(a)21 CFR 11.10(f)

Q6

Is there an audit trail, and does it cover the agent as well as the human?

The two audit-trail clauses, plus retention and archiving — an audit trail that is not retained is not one. For agentic runs the tool path itself is part of the record.

21 CFR 11.10(e)EU GMP Annex 11 §921 CFR 11.10(c)EU GMP Annex 11 §17

Q7

What events invalidate this qualification?

Annex 11 §10 change and configuration management, and §11 periodic evaluation. Part 11 has no direct change-control clause, which is why this question leans on Annex 11.

EU GMP Annex 11 §10EU GMP Annex 11 §11

Q8

What must a qualified person review and sign, and what may the agent never decide?

Authority checks and access control define who may act; signature linking and the Annex 11 signature clause define what a signature binds. Together they are the approval boundary.

21 CFR 11.10(g)EU GMP Annex 11 §1421 CFR 11.7021 CFR 11.10(d)

HOW AN ANSWER IS MARKED

The record contains what this question asks for, and every claim in it resolves to a captured value or a cited clause.

WHAT IT IS NOTIt is not a finding that the answer is good. The question asked whether the thing was shown, not whether a reviewer would accept it.

COUNTS TOWARD A BAND

Some of what the question asks for is in the record and some is not — an input without a checksum, a chain that verifies without an agent path, an evaluation that came back indeterminate.

WHAT IT IS NOTIt is not a passing grade with a caveat. Partial counts for nothing toward a tier, because half-evidenced is the status an assessor reaches for when the work is nearly done and the deadline is not.

COUNTS TOWARD NOTHING

Nobody showed this. It is a finding with a named consequence and a remediation path, and it is the output the assessment exists to produce.

WHAT IT IS NOTIt is not a blank, an error, or a section that failed to load. Something was asked and the answer is that no evidence exists — which is a different sentence from silence.

COUNTS TOWARD NOTHING

The question does not reach this use, and someone wrote down why. The written reason is the requirement, not a courtesy.

WHAT IT IS NOTIt is not a way to clear a question. An N/A with no rationale is treated as not established, because unchecked N/A is the primary route to an inflated tier.

COUNTS TOWARD A BAND

WHAT A COVERAGE BAND COVERS

A band is derived from the answers, never chosen. It is not a grade and there is no passing level — a band says which questions were evidenced for one scope, and nothing about whether a reviewer would accept them.

G0

NO BAND MET

Fewer questions are evidenced than any named band requires. The record names each one, what a reader therefore cannot conclude, and what would close it.

An honest G0 with eight named gaps is a finished deliverable, not a failed run. It is the assessment, and it is the document that scopes the remediation.

G1

Q1Q2Q3Q4Q5Q6Q7Q8

The intended use and the criticality basis are on the record.

It says what the tool is for and how critical that is. It says nothing yet about whether the run can be reproduced or what it was measured against.

G2

Q1Q2Q3Q4Q5Q6Q7Q8

Intended use, criticality, reproducibility, input provenance and pre-run acceptance criteria are all on the record.

This is the band where the criteria were fixed before the result was seen — the property that cannot be added afterwards, and the reason the freeze step is irreversible.

G3

Q1Q2Q3Q4Q5Q6Q7Q8

All eight questions are evidenced for this scope.

It covers one tool, one version and one declared use. It is not a statement about the tool in general, about the next version, or about any other use of the same tool.

WHILE IT STANDS, AND WHEN IT NO LONGER DOES

The record describes the system as it currently runs, within its stated window.

NEXTNothing. The lapse conditions are listed on the record and are watched, not assumed.

A pinned dependency moved — a container digest, a model revision, a reference database, the clause corpus — so the record no longer describes the system that is running.

NEXTRe-run the affected questions. A lapse is the mechanism working: it is what a durable badge cannot do, and it is why the scope on this record can be trusted while it stands.

The record reached the end of the window it was issued for, with nothing else changing.

NEXTRe-issue against the current corpus. Nothing about the prior record is withdrawn.

A later record covers the same scope. The earlier one stays readable and stays verifiable.

NEXTRead the successor. The supersession chain is part of the record and can be walked.

The issuer withdrew the record because something in it was wrong — not because a dependency moved. The stated reason travels with it.

NEXTRead the reason. A revoked record must not be relied on, and a re-issue is a new assessment.

THE GOVERNED SEQUENCE

The order is the product. Step three happens before step four and no later step can move it, which is the one property that cannot be added to a record afterwards.

  1. Declare the context of use. Where this runs, what the output becomes, and what a human does with it. This decides which frameworks bind.
  2. Classify the risk. Software category and criticality, with the operator’s written basis. The tool cannot supply this.
  3. Freeze the acceptance criteria. Thresholds and reasons, hashed. Irreversible: the criteria id is what makes step 5 mean anything.
  4. Execute the invocation. Against pinned versions, through one allowlisted endpoint, inside published input limits.
  5. Capture what happened. Inputs, response hash, observed metrics, missing metrics, timings — into an append-only chain.
  6. Evaluate against the frozen criteria. The verdict is bound to the criteria id. Changing a threshold produces a different id, not a different verdict.
  7. Answer the eight questions. Each status derived from the captured record, each citation resolved against the clause corpus.
  8. Publish, then watch for change. The record states its own lapse conditions. When a pinned dependency moves, the record lapses.

HOW A RUN ENDS

Completed

The invocation returned and was evaluated. The verdict may still be fail or indeterminate — a completed run is not a passing one.

Provider unavailable

The invocation was attempted and produced no result. That is a governed outcome with its own record, not an error page: the verdict is forced to indeterminate and the record says why.

Invalid

The response did not match the contract the adapter declares. Recorded as such rather than parsed hopefully.

WHAT VERIFYING A RECORD DOES NOT ESTABLISH

Carried here in the same words it is carried inside every exported bundle, because a limitation that lives only in marketing copy is a limitation the person holding the file never reads.

  • That the chain was not rewritten wholesale. It is self-anchored: whoever holds the store can produce a consistent chain saying anything. Only external anchoring of the head hash would change that, and SPEQ does not implement it.
  • That the tool behaved correctly, or that a prediction is right. SPEQ records what was run and against what criteria; it does not validate the tool or approve its use.
  • That the operator's declarations are true. Intended use, context of use and criticality are stated by the operator and carried verbatim; SPEQ checks they are present and internally consistent, not that they are accurate.
  • That the regulations named are the only ones that apply. The applicability determination covers the frameworks it names and says so; it is not legal advice.

WHERE THE CLAUSE MAP IS THIN

Four instruments the questions draw on sit in SPEQ’s catalog at instrument level and not at clause level. Naming a part number is a reference, not a pinpoint, so they are listed as gaps rather than cited as though they were.

Present in the standards catalog at instrument level only. Mapping it needs clause-level requirements that do not exist in the corpus yet; citing the part number alone would be a reference, not a pinpoint.

In the standards catalog, no clause-level entries. Its contribution here is the risk-based framing already implemented in the AI risk classifier, not a citable clause.

A copyrighted ISPE guide. SPEQ cites identifiers and does not reproduce text, so a pinpoint clause corpus for it is a licensing question before it is an engineering one.

Scoped to medical-device production and quality-system software. Mapping it onto every AI invocation would be the scope trap the authoring contract warns about.

FREQUENTLY ASKED

Does a Governance Evidence Record mean the AI tool is validated?

No. The record documents what was run, on what, against criteria fixed before the run, with an audit trail — and it states what was not established as prominently as what was. Validation is a determination the regulated party makes under its own quality system. SPEQ supplies structured, cited evidence to support that work; it does not perform it and does not approve the outcome.

Why does the record expire?

Because the system it describes changes. Every record pins the components it depends on — a container digest, a model revision, a reference database, the clause corpus — and states which changes lapse it. A statement that stays true forever about software that changes monthly is not a strong claim; it is an unfalsifiable one. Expiry is the property that lets the scope on a standing record be trusted.

Is a G0 record a failure?

No, and treating it as one is the failure. G0 means fewer questions were evidenced than any named band requires, and the record names each gap, what a reader therefore cannot conclude, and what would close it. That is the assessment deliverable. A system that could not issue G0 would push assessors to inflate statuses, which is exactly what this layer exists to prevent.

What does “not established” mean on a question?

That nobody showed it. It is a finding with a named consequence, not a blank field and not a section that failed to load. It is deliberately distinct from “partial”, which means some of what the question asks for is in the record and some is not. Neither counts toward a coverage band; only one of them describes work already done.

How can someone else verify a record without SPEQ?

By recomputing it. Each record is canonicalised with RFC 8785 JCS and hashed with SHA-256, and the signature is over the recomputed hash rather than over the hash the record claims — so a record edited after issue fails before the signature is even checked. The audit chain links each entry to the previous one by hash. The exported bundle carries the verification steps inside it, because a verifier reads the file they were handed rather than a link to a site that will be redesigned twice.

What does verifying a record NOT prove?

That the chain was not rewritten wholesale. It is self-anchored: whoever holds the store could produce a consistent chain saying anything. Only external anchoring of the head hash would change that and SPEQ does not implement it, so the record says so rather than letting the word “verifiable” imply more than it delivers. Verification also says nothing about whether the tool behaved correctly or whether the operator’s declarations are true.

Which regulations does a record cite?

Whichever ones the declared context of use brings into scope, and the record says which. 21 CFR Part 11 reaches records a predicate rule requires, not electronic records in general, so a run whose output informs no such record cites Part 11 as reference rather than as binding. Over-citation reads as conservative and is not: it asserts an obligation the operator may not have, and it teaches readers to discount the citations.

Is SPEQ affiliated with the tool vendors it builds adapters for?

No. An adapter is SPEQ’s description of what can be captured about one tool family, built against publicly documented interfaces. It carries no affiliation, endorsement or approval from the vendor, and an adapter existing is not a statement that the tool is suitable for any regulated use.