OPENGAMP 5FDA/EMA GAIP (2026)ISO/IEC 42001

AI Evaluation Score & Release Gate

Score an AI system’s evaluation results against SPEQ’s GxP benchmark dimensions — requirement, citation, source-authority, jurisdiction, hallucination, applicability, and more — and get a deterministic release verdict. Critical-dimension failures (a fabricated citation, the wrong jurisdiction) block release regardless of the aggregate, the way the assurance rule intends: the benchmark is SPEQ’s and the model never grades itself. A scoping aid, not a validated assurance system.

OUTPUT

Overall score + pass / conditional / fail verdict

TIME

~8 min

Limitations — read before you rely on this

  • This is a scoping aid, not a validated system, and not a substitute for your own qualified evaluation harness. Reproduce the scoring in a system you control and retain the case-level results; a copied verdict is not evidence.
  • It scores the rates you enter — it does not run the evaluation. The integrity of the verdict depends entirely on the suite behind the rates: its size, its representativeness, and the ground truth it is graded against, all of which are yours to defend.
  • The critical floor and overall target are SPEQ judgement, not FDA/EMA thresholds. A different, defensible benchmark could set them elsewhere; treat a near-boundary verdict as a prompt to look at the case-level failures, not a certificate.
  • It collapses a suite to one aggregate rate per dimension. Real release decisions should inspect the individual critical-dimension failures case by case, because one fabricated citation on a safety-critical question matters more than the percentage suggests.

WHAT THIS CALCULATES

A deterministic release verdict for an AI system from its evaluation results: it rolls each of SPEQ’s GxP benchmark dimensions (requirement, citation, source-authority, jurisdiction, hallucination, applicability, and more) up to an overall score, then applies the assurance rule that a failure on any critical dimension blocks release regardless of the aggregate. It turns "the eval looked good" into pass / conditional / fail, with the reason attached — and the benchmark is SPEQ’s, so the model never grades itself.

THE METHOD

overall = mean( dimension_rate ); verdict = fail if any critical_dim < 50%, else conditional if overall < target, else pass
dimension_rate
the pass rate for one benchmark dimension across the suite — the share of cases the system got right on that dimension (0–100%)
overall
the equal-weighted mean of the evaluated dimensions’ pass rates
critical_dim
a dimension whose failure makes an answer unsafe to rely on in a regulated setting — requirement, citation, source-authority, jurisdiction, and freedom from hallucination
target
the overall score a suite must clear (80%) to read as a clean pass rather than conditionally acceptable
verdict
the deterministic outcome: pass, conditional (remediate the weak dimensions), or fail (a critical failure or an empty suite)

The 50% critical floor and the 80% overall target are SPEQ judgement expressing the epic’s "critical eval fails → not passing" rule, not regulatory constants. Scoring is deterministic and monotonic: raising any dimension’s rate never lowers the overall nor worsens the verdict.

THE INPUTS, AND WHAT THEY MEAN

Per-dimension pass rate
For each benchmark dimension, the fraction of your evaluation cases the system answered correctly on that dimension. This is the outcome of running your suite against SPEQ’s rubric — not the model’s own opinion of its answers, which is exactly the input the assurance model excludes.
Critical vs non-critical (fixed)
Which dimensions are release-blocking is set by SPEQ’s benchmark, not by the user: a fabricated citation, a superseded source cited as current, or the wrong jurisdiction’s law is a hard failure however strong the rest of the answer is. You score the rate; the criticality is the rubric.
[ AI EVALUATION · SPEQ BENCHMARK ]

Turn evaluation results into a release verdict — deterministically.

Score an AI system against SPEQ’s GxP benchmark dimensions and get a pass / conditional / fail verdict the assurance way: the benchmark is SPEQ’s, the grading is deterministic, and the model never grades itself. A failure on any critical dimension — a fabricated citation, a superseded source, the wrong jurisdiction — blocks release regardless of the overall score.

Enter each dimension’s pass rate across your evaluation suite (the share of cases the system got right on that dimension). Critical dimensions are marked — if one passes less than half the time, the suite fails.

RELEASE VERDICT
PASS
Overall 90% · target 80%.
WHY

No critical failures and a 90% overall score at or above the 80% target.

PROFESSIONAL EXPORT

HOW TO READ THE OUTPUT

  • A high overall score with a critical failure is still a fail. That is the whole point — an AI answer that invents a citation or applies the wrong jurisdiction is unsafe to act on no matter how polished the rest is, so the gate refuses to average that away.
  • A "conditional" verdict is an instruction, not a grade: no critical failure, but the aggregate is below target, so the honest next step is to remediate the weakest dimensions and re-run, not to ship.
  • Read abstention as a strength. A system that declines to answer when its grounded sources do not support one should score well on that dimension; a benchmark that punishes appropriate abstention rewards confident guessing.
  • The verdict is only as good as the ground truth behind the rates. The dimensions come from SPEQ’s benchmark, but the pass rates come from your suite — a thin or unrepresentative suite produces a confident verdict on very little, which is itself worth noticing.

WORKED EXAMPLE

A RAG regulatory copilot is evaluated on 40 GxP questions. It scores well on most dimensions but fabricates a clause citation on 30% of cases (70% citation pass rate) — below the 50% floor? No: check the rule against the critical rate.

Citation accuracy (critical)
45%
Source authority (critical)
95%
Jurisdiction (critical)
100%
Other dimensions
85–100%

RESULT

Overall ≈ 90%, but citation accuracy 45% < 50% critical floor → verdict: FAIL (release blocked)

The suite would look like a 90% pass on the headline number, and it is a fail: a critical dimension — citation accuracy — dropped below the floor, so the deterministic gate blocks release regardless of the aggregate. The remediation is specific: fix the citation grounding, re-run, and the same 90% becomes a pass once the critical dimension clears the floor.

REGULATORY BASIS

FDA/EMA — Guiding Principles of Good AI Practice in Drug Development (Jan 2026)
Frames the model-independent, benchmark-owned evaluation of AI in regulated development that this scorer operationalises; see SPEQ’s Good AI Practice page.
ISPE GAMP 5 (2nd ed.) — A Risk-Based Approach to Compliant GxP Computerized Systems
Supplies the risk-based assurance framing under which an AI system’s evaluation results can become validation evidence.
ISO/IEC 42001:2023 — AI management systems
The management-system expectation for defining, measuring, and governing AI performance that a documented benchmark supports.
PROFESSIONAL · WORKED SCENARIOS · SPEQ SYNTHESIS

See this tool applied to real cases

CHECKING ACCESS

Checking your Professional access…