Source Data, eSource & Direct Data Capture
A surprising amount of confusion in clinical data management dissolves once two distinctions are clear: source data versus source documents, and the family of ways source data is captured electronically. Source data is the information; source documents are the media that hold it. eSource is the umbrella for capturing that information electronically from the start, and direct data capture is the subset where the electronic case report form is itself the source — with consequences that reach all the way to whether verification is even possible. This page settles those distinctions; the monitoring activities that act on the data are covered in the [SDV vs SDR](/topics/source-data-verification) explainer. The FDA’s 2013 Electronic Source Data guidance and 21 CFR Part 11 frame the expectations.
What an explainer is not
A topic explainer is SPEQ’s synthesis of what a practice involves, cited to the standards that govern it. It does not reproduce their text, and it does not determine which of them apply to your product or process.
[ POSITION IN THE FRAMEWORK ]
7 DIMENSIONS · 20 LINKSSource data is the information; source documents are the containers. Under direct data capture the eCRF is itself the source, which makes source-data verification conceptually impossible and shifts the control to the Part 11 audit trail.
06 · QUALITY MATURITY — SOURCE DATA, ESOURCE & DIRECT DATA CAPTURE, REACTIVE TO ADAPTIVE
Source and source documents are used interchangeably; DDC fields have no prior source and no designation, so provenance cannot be reconstructed.
eSource is adopted but treated as one technology; which fields are DDC is not prospectively identified, leaving verification realities unclear.
A prospective source-data identification list states, per element, where the source is, so monitors know what is verified versus reviewed.
Part 11 audit trails, access controls, and edit history are the integrity mechanism where the eCRF is source; provenance is traceable end to end.
Mixed-capture trials are auditable by design; source map, audit trail, and review layers compose so any data point's origin is answerable.
SPEQ’s shared five-stage progression, labelled synthesis — not the FDA QMM rating scale. Where does your organization sit? Score your quality system →
07 · REGULATORY & EVIDENCE
GOVERNING STANDARDS · 3
Derived from the 3 standards SPEQ maps to this subject, across 2 regulatory bodies: FDA, ICH.
RECORDS & OBJECTIVE EVIDENCE
- A prospective source data identification list naming the source for each data element
- Designation of which fields are direct data capture (eCRF as source)
- Part 11 audit trails, access controls, and edit history for eSource systems
- Data-flow traceability from original observation to CRF to analysis dataset
- Validation records for the EDC / eCOA / eConsent systems in use
COMMON INSPECTION FINDINGS
- Source undesignated in a DDC-heavy trial, so provenance cannot be reconstructed
- Source data and source documents conflated
- 'We use eSource' with no modality or control specified
- SDV expected on DDC fields that have no prior source
- eSource systems without Part 11 audit trails or access controls
Source data vs source documents — information vs container
The underlying distinction is content versus medium. **Source data** is the information itself — the original observation, result, or record of a clinical finding. **Source documents** are the physical or electronic containers that hold it: the medical chart, the lab printout, the ECG trace, the pharmacy log. One is the data; the other is where the data lives. Conflating them ("the source document is the source data") is the error beneath much of the confusion, because the two behave differently — the same source datum can exist in more than one document, and the question of *which* record is the source of truth has to be answered deliberately.
This is not pedantry: regulators expect a trial to be able to say, for any data point, where it originated and how it flowed to the case report form. That traceability — original observation → source document → CRF → analysis dataset — is the spine of data integrity in a trial, and it is only expressible if you keep straight what the data is and which container is authoritative for it.
eSource — the umbrella, not one thing
eSource means source data captured initially in electronic form, and it is an **umbrella term** covering several modalities: extraction from electronic health records, electronic clinical outcome assessments (eCOA, including patient-reported ePRO), electronic informed consent (eConsent), and direct data capture. What unites them is that the *first* durable record of the data is electronic rather than paper transcribed later. The FDA’s 2013 guidance, Electronic Source Data in Clinical Investigations, is the anchor: it promotes electronic capture and sets expectations around identifying authorised data originators, data-element identifiers for the audit trail, and how source data flows into the eCRF.
Because eSource is an umbrella, using it as if it named a single technology is a category error — "we use eSource" says almost nothing without specifying which modality. Each has its own control questions: EHR extraction raises provenance and transformation concerns; eCOA raises questions of who the authorised originator is; eConsent raises Part 11 signature and version-control questions. And because these are electronic records and signatures in an FDA-regulated study, 21 CFR Part 11 expectations — audit trails, access controls, attributability — apply across them.
Direct data capture — when the eCRF is the source
**Direct data capture (DDC)** is the specific subset of eSource where data is entered straight into the electronic case report form with no prior source document — the eCRF *is* the source. A coordinator recording a vital sign directly into the study system at the point of measurement, with no paper worksheet behind it, is doing DDC. This is efficient and increasingly common, and it has a consequence that surprises people the first time they meet it: **source data verification is conceptually impossible for a DDC field.** SDV compares the CRF to a prior source — but under DDC there is no prior source; the CRF entry is the original record. There is nothing to verify it *against*.
That does not mean DDC fields escape oversight — it means the controls shift. Because the eCRF is the source, its Part 11 audit trail, its access controls, and its edit history become the integrity mechanism, and source data *review* and centralised monitoring replace the impossible verification. It also makes one control non-negotiable: the trial must **prospectively designate which fields are DDC** — a source data identification list naming where the eCRF is the source — so that everyone, including an inspector, knows which data has a prior source to verify against and which does not. Without that designation, the arrangement is uninspectable; with it, DDC is a controlled, defensible source.
Why the designation is the whole control
The single most important practical takeaway is that the *identification of source* must be prospective and documented. A trial mixing paper source, EHR-extracted data, eCOA, and direct data capture has different verification realities for different fields — some have a prior source to check against, some do not — and the only way that is inspectable is a source data identification plan written before data collection, stating for each data element where the source is. This is the document that turns a technically complex, mixed-capture trial into an auditable one.
Get this right and the rest follows: monitors know what they can verify and what they must instead review; the audit trail is the integrity control where the eCRF is the source; and the trial can answer, for any data point, the regulator’s question of where it came from. Get it wrong — leave source undesignated in a DDC-heavy trial — and you have data whose provenance cannot be reconstructed, which is a data-integrity finding regardless of how good the underlying systems are.
FREQUENTLY ASKED
What is the difference between source data and source documents?
Source data is the information — the original observation, result, or clinical finding. Source documents are the containers that hold it: the medical chart, lab printout, ECG trace, pharmacy log. One is the data; the other is where it lives. The same source datum can exist in more than one document, so which record is the source of truth must be decided deliberately — the basis of data traceability in a trial.
What is eSource?
An umbrella term for source data captured initially in electronic form, covering several modalities: electronic health record extraction, electronic clinical outcome assessments (eCOA/ePRO), electronic informed consent (eConsent), and direct data capture. "We use eSource" says little without specifying which modality, since each has distinct control questions. The FDA’s 2013 Electronic Source Data guidance and 21 CFR Part 11 frame the expectations.
Why is source data verification impossible under direct data capture?
Because direct data capture (DDC) enters data straight into the eCRF with no prior source document — the eCRF is the source. SDV compares the CRF to a prior source, but under DDC there is no prior source to compare against. Oversight shifts to the eCRF’s Part 11 audit trail and access controls, plus source data review and centralised monitoring. The trial must prospectively designate which fields are DDC so it is clear which data has a source to verify against.
What makes a mixed-capture trial inspectable?
A prospective, documented source data identification list stating, for each data element, where the source is — paper, EHR extraction, eCOA, or direct data capture. This is the control that lets monitors know what can be verified versus reviewed, makes the audit trail the integrity mechanism where the eCRF is source, and lets the trial answer where any data point came from. Leaving source undesignated in a DDC-heavy trial is a data-integrity finding.