EDC
Electronic Data Capture
An EDC system is the clinical database of a trial: the electronic case report forms on which sites enter subject data, the edit checks that challenge implausible entries as they are made, the query workflow through which data management pursues discrepancies, and the locked dataset that statistics ultimately analyses. Every conclusion the trial reaches — efficacy, safety, the label claim — rests on what this system captured and how defensibly it was cleaned. That is why the audit trail requirement lands here with full force: ICH E6(R3) expects any change or correction to trial data to be traceable, to not obscure the original entry, and to be explained where explanation is needed.
What this page does not claim
A system class is not a product. SPEQ describes what a CTMS or a LIMS is; the vendor directory at /tools lists the products that implement one, and a GAMP category is a property of an implementation, not of a class.
What a EDC actually is
An EDC system is the clinical database of a trial: the electronic case report forms on which sites enter subject data, the edit checks that challenge implausible entries as they are made, the query workflow through which data management pursues discrepancies, and the locked dataset that statistics ultimately analyses. Every conclusion the trial reaches — efficacy, safety, the label claim — rests on what this system captured and how defensibly it was cleaned. That is why the audit trail requirement lands here with full force: ICH E6(R3) expects any change or correction to trial data to be traceable, to not obscure the original entry, and to be explained where explanation is needed.
The distinction that matters most is between the entered record and the source. In the classic model, source data originates at the site — medical records, instrument printouts, worksheets — and the eCRF is a transcription of it, which is what source data verification exists to check. Increasingly, data also enters the eCRF as eSource, captured electronically at first observation with no paper predecessor, or arrives by validated transfer from central laboratories, eCOA platforms, and the IRT. Each route changes what "verification" means, and a trial's data-governance documentation must say, per data element, where source lives.
The system's life runs in two registers. The platform is validated once and maintained under change control; the study build — the eCRF designs, edit-check specifications, derivations, and integrations configured for a specific protocol — is created fresh for every trial and tested against the protocol it serves. Most defects that matter are build defects: an edit check that fires on the wrong branch, a visit structure that cannot represent a real subject's path, a unit conversion that corrupts silently. Database lock is the ceremony that ends the second register: after lock, the dataset is frozen, and any unlock is a controlled, documented exception.
EDC is unambiguously a 21 CFR Part 11 system — site investigators sign eCRF casebooks electronically, and those signatures carry the investigator's attestation of the data. ICH E6(R3) is what distributes the validation burden here: it requires computerised systems to be fit for purpose and controlled proportionately to the risk they carry to participant safety and data reliability, which concentrates rigour on the functions bearing on data integrity and leaves lighter-touch evidence elsewhere. (FDA's Computer Software Assurance guidance is often invoked for the same philosophy, but its stated scope is software used in medical-device production and quality systems, not clinical trial systems.) In the EU, trials run under Regulation (EU) 536/2014 with GCP inspection reach into the sponsor's computerised systems, and the MHRA's data-integrity expectations apply to the clinical database exactly as they do to a laboratory one: ALCOA+ is the shape of every argument about whether the data can be believed.
WHERE THE BOUNDARY ACTUALLY SITS
Not the source in the default model. Source data originates at the site or in the connected system that first captured it; the eCRF is the entered record unless the trial explicitly designates direct entry as eSource.
Not the outcome-assessment platform. Patient- and clinician-reported outcomes are captured in the eCOA system, which holds their source and audit trail; the EDC receives them by validated integration.
eCOA / ePRO owns it →Not the safety case processor. An SAE recorded on an eCRF triggers reporting obligations, but the individual case safety report is built, assessed, and submitted from the pharmacovigilance database, and the two are reconciled.
Safety / PV Database owns it →Not the analysis environment. Locked data is extracted to the statistical environment; derivations for analysis, the analysis datasets, and the outputs of ICH E9-governed statistical work live outside the EDC.
WHAT IT HOLDS, AND WHAT CROSSES ITS BOUNDARY
CORE RECORDS
- eCRF data for every subject, with the full audit trail of entries, changes, and reasons for change
- Edit-check specifications and their execution history
- Query records — raised, answered, and closed — with the data changes they produced
- Study build documentation: annotated CRFs, configuration specifications, and build testing evidence
- Investigator electronic signatures on casebooks, and what each signature attested
- Database lock and unlock records, with the authorisation and rationale for any unlock
- User access history — who held which role at which site, and when access was granted and retired
DATA FLOWS OUT
SAE data reconciled field-by-field against the safety case, with discrepancies queried back to the site
Enrolment, visit, and data-entry status metrics that feed monitoring triggers and sponsor oversight
Archived casebooks, data-management deliverables, and build documentation filed as essential records at closeout
HOW THIS CLASS IS USUALLY VALIDATED
- SPEQ synthesis: EDC splits cleanly into a GAMP 5 Second Edition (2022) Category 4 platform — validated once, maintained under change control — and a per-study build whose eCRFs, edit checks, and integrations are configuration verified against each protocol. The category attaches to the implementation, not the product, and the per-study layer is where the risk concentrates.
- Study-build testing is protocol-driven: edit checks exercised against passing and failing data, branching and visit structures walked with realistic subject paths, and integrations proven with round-trip data — because a build defect ships wrong data quietly for the life of the trial.
- Part 11 verification is direct, not inherited: investigator signature linkage to the casebook, the impossibility of altering signed data without invalidating the signature, and an audit trail that no site or sponsor role can suppress.
- Under a CSA-leaned approach, the functions that decide data integrity — audit trail, signatures, lock, randomisation-blind protections on integrated data — get scripted evidence; administrative conveniences do not need the same ceremony.
SPEQ synthesis, not a rating. This is SPEQ’s reading of how this system class is commonly approached, offered to help you scope your own work. A GAMP category is a property of a specific implementation, not of a product class, and one deployment routinely spans several. It is not a classification service and does not replace your own documented risk assessment.
- Stage 1 · Reactive
The build is tested informally and defects surface in production as mid-study migrations. Queries pile up unworked, the audit trail is never reviewed, and cleaning happens in a rush before lock, where late discoveries force choices between the timeline and the data.
- Stage 2 · Defined
Build, test, and release of each study database follow a defined process with documented evidence. Query and data-review conventions exist, but metrics are compiled manually, review effort is spread evenly rather than by risk, and mid-study amendments still strain the change process.
- Stage 3 · Controlled
Edit checks, review listings, and reconciliations run to a documented data-management plan, query ageing and site data-quality metrics are tracked continuously, and protocol amendments flow through controlled build changes with regression evidence. Lock is a scheduled event, not a crisis.
- Stage 4 · Predictive
Data review is risk-proportionate and centralised: statistical and pattern-based checks surface anomalous sites and implausible data that field monitoring would miss, SDV is targeted where source risk is real, and audit-trail review is a designed activity with documented scope.
- Stage 5 · Adaptive
Data quality is engineered upstream — standards-based build libraries, reusable validated checks, and eSource-first design remove transcription instead of policing it. The organisation measures error rates at the point of origin and improves the capture design, not just the cleaning.
SPEQ’s shared five-stage progression, labelled synthesis. It is not the FDA QMM rating scale and not the scored maturity-assessment domains — assess your quality system for those.
WHAT AN INSPECTION PROBES, AND WHERE IT GOES WRONG
INSPECTION SIGNALS
- Whether the audit trail survives scrutiny — original values preserved, changes attributed and explained, and no role able to switch the trail off.
- Whether data changes after investigator signature invalidated and re-required the signature, and whether the investigator actually reviewed what was attested.
- The query record as a portrait of site data quality — volumes, ageing, and whether repeated failures at a site fed back into monitoring.
- Database lock discipline: who authorised lock, what changed between soft and hard lock, and the justification and control around any unlock.
- Whether user access matched delegation — entries by staff not on the delegation log, or access surviving a site's closure, undermine attributability.
COMMON RISKS
- Build defects that pass superficial testing — a mis-scoped edit check or broken visit branch — and quietly distort data until a mid-study migration is forced.
- Reconciliation between EDC and the safety database left until lock, surfacing SAE discrepancies at the worst possible moment.
- Audit-trail review committed to in SOPs but never performed against actual records, discovered by the inspector rather than the sponsor.
- Integration mappings — labs, eCOA, IRT — verified once and never revisited after upstream format changes, corrupting data without raising an error.
- Shared or generic site logins that make eCRF entries unattributable, collapsing the first letter of ALCOA.
WHO WORKS IN IT, AND WHERE IT IS SHAPED
ROLES
- Clinical data manager
- Clinical database programmer / study builder
- Site study coordinator (data entry)
- Clinical Research Associate (source data verification)
- Biostatistician
- CSV analyst
DELIVERY-LIFECYCLE PHASES
[ POSITION IN THE FRAMEWORK ]
6 OF 7 DIMENSIONS · 23 LINKSThe clinical database — where subject data is captured on eCRFs, edit-checked, queried, cleaned, and locked for analysis; every efficacy and safety conclusion rests on what it captured and how defensibly it was cleaned.
06 · QUALITY MATURITY — EDC, REACTIVE TO ADAPTIVE
The build is tested informally and defects surface in production as mid-study migrations. Queries pile up unworked, the audit trail is never reviewed, and cleaning happens in a rush before lock, where late discoveries force choices between the timeline and the data.
Build, test, and release of each study database follow a defined process with documented evidence. Query and data-review conventions exist, but metrics are compiled manually, review effort is spread evenly rather than by risk, and mid-study amendments still strain the change process.
Edit checks, review listings, and reconciliations run to a documented data-management plan, query ageing and site data-quality metrics are tracked continuously, and protocol amendments flow through controlled build changes with regression evidence. Lock is a scheduled event, not a crisis.
Data review is risk-proportionate and centralised: statistical and pattern-based checks surface anomalous sites and implausible data that field monitoring would miss, SDV is targeted where source risk is real, and audit-trail review is a designed activity with documented scope.
Data quality is engineered upstream — standards-based build libraries, reusable validated checks, and eSource-first design remove transcription instead of policing it. The organisation measures error rates at the point of origin and improves the capture design, not just the cleaning.
SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →
07 · REGULATORY & EVIDENCE
GOVERNING STANDARDS · 8
Derived from the 8 standards SPEQ maps to this subject, across 5 regulatory bodies: FDA, ISPE, ICH, EMA, MHRA.
RECORDS & OBJECTIVE EVIDENCE
- eCRF data for every subject, with the full audit trail of entries, changes, and reasons for change
- Edit-check specifications and their execution history
- Query records — raised, answered, and closed — with the data changes they produced
- Investigator electronic signatures on casebooks, and what each signature attested
- Database lock and unlock records, with the authorisation and rationale for any unlock
COMMON INSPECTION FINDINGS
- An audit trail that can be switched off, or data changes not attributed and explained
- Data changed after investigator signature without invalidating and re-requiring the signature
- Build defects — a mis-scoped edit check or broken visit branch — distorting data until a mid-study migration
- SAE reconciliation between the EDC and the safety database left until lock
- Shared or generic site logins that make eCRF entries unattributable
Choosing, validating, and living with EDC
SPEQ curates the software directory and does not endorse, certify, or rank any vendor. Listing is not a recommendation, and this catalog describes the system class, not the product.
FREQUENTLY ASKED
Is an EDC system a 21 CFR Part 11 system?
Yes, about as centrally as any system in clinical research. The eCRF data it holds are electronic records required under the predicate rules governing the trial, and the investigator's electronic signature on the casebook is a regulatory signature attesting to the data. Part 11 therefore requires secure, computer-generated, time-stamped audit trails; signatures linked to their records so they cannot be excised or transferred; and access limited to authorised individuals. In practice, inspectors probe the audit trail and the signature-invalidation behaviour on data change more than any other EDC function.
What is the difference between EDC data and source data?
Source data is the first durable capture of an observation; EDC data is what was entered into the trial database. In the traditional model those differ — the site records into medical notes or worksheets, then transcribes into the eCRF, and source data verification checks the transcription. When the trial designates direct eCRF entry as eSource, or data arrives electronically from eCOA, central labs, or devices, the first capture and the database record converge and transcription checking becomes meaningless. What matters is that the protocol's data-governance documentation states, per data element, where source resides.
Can data be changed after database lock?
Only as a controlled exception. Lock marks the point at which cleaning is complete and the dataset is frozen for analysis; after it, edit rights are removed and the data is what the statistics will be run on. If a genuine error is discovered post-lock — often during analysis or reconciliation — the database can be unlocked, but the unlock is authorised at a senior level, scoped to the specific correction, documented with rationale, and closed with a re-lock. An unlock that looks routine is a finding: it suggests lock was declared before cleaning had actually finished.
Does every study need its own validation of the EDC system?
Every study needs its build verified; no study needs the platform revalidated. The platform — the software that renders forms, runs checks, and writes the audit trail — is validated once as configured commercial software and maintained under change control. Each protocol then gets a study-specific build: eCRF designs, edit checks, derivations, visit structures, and integrations. That build is tested against the protocol before go-live, and amended under change control mid-study. The recurring failure is treating build testing as a formality because "the system is validated" — the platform being sound says nothing about whether this study's checks are.