Sampling Plans, Specifications & Standards
What gets tested and against what: sampling plans and their statistical basis, specifications and their justification, reference standards, and retention samples. A result describes the sample, and the sample only describes the batch if the sampling plan makes it representative. Sampling is the step where the strength of every downstream conclusion is actually set, and it receives the least scrutiny.
What an explainer is not
A topic explainer is SPEQ’s synthesis of what a practice involves, cited to the standards that govern it. It does not reproduce their text, and it does not determine which of them apply to your product or process.
[ POSITION IN THE FRAMEWORK ]
7 DIMENSIONS · 22 LINKSA specification is a claim about a batch made from a fraction of it, so the sampling plan is doing the load-bearing work — and a statistically sound plan drawn from an unrepresentative location proves nothing.
06 · QUALITY MATURITY — SAMPLING PLANS, SPECIFICATIONS & STANDARDS, REACTIVE TO ADAPTIVE
Samples are taken as they always have been. The plan is in the procedure and its origin is unknown.
Plans state quantity and location and reference a compendial requirement, with the locations chosen for accessibility.
The plan derives from what the sample must represent — where variability actually occurs, at what stage — and specification limits derive from process capability and clinical relevance rather than from what has been achieved.
Sampling is reviewed when the process, scale or material changes, and the plan is treated as part of the control strategy rather than as a laboratory procedure.
Specifications are set where they discriminate — narrow where the attribute matters, wide where it does not — and the sampling supports that discrimination.
SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →
07 · REGULATORY & EVIDENCE
GOVERNING STANDARDS · 5
Derived from the 5 standards SPEQ maps to this subject, across 5 regulatory bodies: FDA, EMA, EDQM, ICH, USP.
RECORDS & OBJECTIVE EVIDENCE
- Sampling plans with the rationale for location, quantity and stage
- The link between sampling design and where variability is known to occur
- Specification justification, including capability and clinically relevant data
- Sampling tools and techniques, and their qualification where they affect the result
- Review records where a process or scale change prompted resampling design
COMMON INSPECTION FINDINGS
- Sample locations chosen for access rather than for representativeness
- Specification limits set from manufacturing history with no clinical or capability basis
- A sampling plan unchanged through a scale or equipment change that moved the variability
- Sample size that cannot detect the variability the specification is meant to control
- Sampling technique introducing bias, unassessed
The plan carries a statistical claim, stated or not
Every sampling plan implies a confidence statement about the batch. Taking a fixed number of units, or the square root of the container count plus one, is a convention rather than a derivation, and the convention encodes an assumption about how variability is distributed through the batch. Where the material is genuinely homogeneous that assumption is safe; where segregation or stratification is possible it may not be.
The plans worth deriving rather than inheriting are the ones supporting a decision about uniformity: blend and content uniformity, stratified sampling across a compression run, or any case where the question is whether a property varies across the batch rather than what its average is. USP <1220> frames the analytical procedure lifecycle around the decision the result must support, and sampling is the first step in that chain.
Specifications justified against the decision they serve
ICH Q6A frames specification-setting as choosing which tests and acceptance criteria are necessary to assure quality, justified from development data and manufacturing experience. Two failure directions follow. Criteria set too tight generate out-of-specification results with no product-quality meaning, consume investigation capacity and eventually train people to expect the investigation to conclude nothing. Criteria set too wide fail to detect the change they exist to detect.
The useful question for each criterion is what decision it supports and what would be missed if it were absent. Criteria that cannot answer that are usually inherited — from a monograph, from a similar product, or from a development-stage limit that was never revisited once the process was understood.
Reference standards are a silent source of drift
Every quantitative result is relative to a reference standard, and the standard has a characterisation, an expiry, storage conditions and a qualification chain of its own. A secondary standard qualified against a primary, used past its intended life or stored incorrectly, shifts every result generated against it — consistently, in one direction, and invisibly.
The characteristic failure is that this is discovered when a new standard lot is introduced and results step. At that point the question is which of the two periods was correct, and answering it requires records for both. Pharmacopoeial standards and their qualification chain deserve the same lifecycle attention as instruments, and they usually receive considerably less.
SPEQ interpretation — the retention sample is the only physical evidence left
Retention samples are stored as a regulatory obligation and are treated as archive. They are, in fact, the only physical evidence of what a batch actually was — and the only way to answer a question that arises years later, when a complaint, a stability signal or an investigation asks something nobody thought to test at release.
That reframes two decisions usually made administratively: what is retained, and under what conditions. A retention sample stored in conditions that let it degrade differently from the marketed product cannot answer the question it exists for, and a retention that omitted a presentation or a component means the question about that one is unanswerable. Both are cheap to get right and impossible to fix retrospectively.
FREQUENTLY ASKED
Is a conventional sampling plan good enough?
For homogeneous material, usually. Conventions such as a fixed count or square-root-plus-one encode an assumption about how variability distributes through the batch, and where segregation or stratification is possible that assumption may not hold. Plans supporting a uniformity decision are worth deriving rather than inheriting.
How tight should an acceptance criterion be?
Tight enough to detect what it exists to detect and no tighter. Criteria set too tight generate OOS results with no product-quality meaning and train people to expect investigations to conclude nothing; too wide and they miss the change. The test for each criterion is what decision it supports and what would be missed without it.
Why are reference standards a source of drift?
Because every quantitative result is relative to one. A secondary standard used past its intended life or stored incorrectly shifts results consistently, in one direction, and invisibly — usually discovered when a new lot is introduced and results step, at which point answering which period was correct needs records for both.
Why do retention samples matter beyond compliance?
They are the only physical evidence of what a batch actually was, and the only way to answer a question arising years later that nobody thought to test at release. Storage conditions that let them degrade differently from marketed product, or a retention that omitted a presentation, make that question permanently unanswerable.