Your model works. What now?
For clinician-researchers, imaging and data scientists, software teams and the people who end up owning quality — translating an imaging result toward something a patient could be affected by. Answer six questions and you get where you are, what is unresolved, what to start keeping today, and where the agency’s own answer lives.
Nothing here determines device status, a regulatory pathway, a submission type, or whether a model is clinically adequate. Membership of a workflow family is not a regulatory determination. Whether a particular function is a regulated device turns on its intended use, its user, its output, the action it influences and the claim made for it — not on the family it belongs to.
START HERE — SIX QUESTIONS
No account, nothing saved, nothing sent. The answers stay in this browser tab.
You have not settled whether this is a research tool or a product. That is a normal place to be and an expensive place to stay, because evidence gathered now is only valid for whichever answer you land on.
- Whether this is a research tool or a product you intend people to use in care. Everything below reads differently depending on the answer.
- What the model actually produces, written down in one sentence a colleague outside the team would understand.
- Whether performance holds outside the site the data came from. One site is not a limitation to hide; it is a scope to state.
- Whether any data has been held back untouched. Once a set has informed a choice — a threshold, an architecture, a stopping point — it is no longer independent of the result.
- Every version of the sentence describing what the model is for, and what changed between them
- Which data came from where, on what equipment, and under what agreement
- Which exact model produced each result you are relying on, and what it was trained on
- What you decided not to do, and why — architecture, threshold, population, exclusions
Saves your six answers, not the sentence you typed. What the model produces is your description of your own product and is never sent anywhere.
The agency’s own words. SPEQ prepares the question and does not answer it.
Open stage 01 — Context of use first.
THE TWELVE STAGES
The same twelve stages apply to every AI-enabled clinical workflow. What changes between families is the evidence emphasis, not the sequence.
Context of use
Who uses this, on whom, where, and what do they do next?
Almost everything downstream follows from this rather than from the technology. Regulators call it the context of use, and evidence gathered against one context does not transfer to another.
WHAT YOU WILL BE ASKED
- Who is in front of the screen when your output appears?
- What were they about to do, and what do they do differently because of it?
- What kind of patient is this, and who is excluded?
- Is this a scanner in a large teaching hospital, a community clinic, or both?
START KEEPING NOW
The description of the setting, written down and dated, before it drifts between the deck and the protocol
WHO OWNS THIS
- Clinician-researcher
- You own this stage. You are the only person who can say what the clinician was about to do without the model.
- Imaging / data scientist
- This decides which data is representative. A context settled after collection makes the collection a coincidence.
- Software team
- This is where the display, the timing and the integration constraints come from.
- Quality and regulatory
- This is the sentence every other document has to agree with. Control it from the day it exists.
What you will claim
What are you willing to say it does, in writing, to someone who will hold you to it?
The claim is what evidence is gathered against and what promotional material is read against. A capability nobody claims needs no evidence; a claim nobody can evidence is the expensive kind.
WHAT YOU WILL BE ASKED
- Say what it does in one sentence, without the words "AI" or "platform".
- What would make that sentence false?
- Does your website say something stronger than that sentence?
- Does your grant application say something different again?
START KEEPING NOW
Every version of the claim, and what changed
Marketing and pitch material as issued
WHO OWNS THIS
- Clinician-researcher
- Resist the strongest defensible claim. The narrowest claim you can evidence is the cheapest one to hold.
- Imaging / data scientist
- The claim sets the metric. A claim about triage and a claim about diagnosis are not the same evaluation.
- Software team
- The claim constrains what the interface may say, including tooltips and result labels.
- Quality and regulatory
- Reconcile the claim against every artefact that describes the product. The mismatch is found easily.
Your regulatory position, as a hypothesis
What do you believe about how this is regulated, and what would change your mind?
A pre-submission meeting rewards reasoned positions with named uncertainties. It treats a confident conclusion with nothing behind it as the thing to probe. IMDRF N10 and N12 give the vocabulary for stating the position.
WHAT YOU WILL BE ASKED
- What do you think this is, and why?
- What is the strongest argument that you are wrong?
- Which country do you intend to reach first, and does that change the answer?
- What have you assumed because someone said it in a meeting?
START KEEPING NOW
The position, its reasoning, and the date — so a later change is visible as a change
WHO OWNS THIS
- Clinician-researcher
- Your clinical framing is an input here, not the answer. The same clinical utility can sit under different frameworks.
- Imaging / data scientist
- The position decides what evidence is expected of the model, so it is worth knowing before the study is designed.
- Software team
- It also decides what the software lifecycle records must look like.
- Quality and regulatory
- You own the memo. Name the open questions — hiding them is what diligence finds.
Where the data came from
For every image you trained on, can you say where it came from and under what agreement?
Provenance is the class of record that cannot be reconstructed. A dataset assembled from a shared drive with no origin is not evidence, whatever the model does with it.
WHAT YOU WILL BE ASKED
- Which sites, which scanners, which years?
- What agreement covered the transfer, and does it permit commercial use?
- Was anything de-identified, and by whom?
- Could you produce that answer for an auditor next year, from records rather than memory?
START KEEPING NOW
Source, agreement and equipment for each dataset, recorded now rather than reconstructed
WHO OWNS THIS
- Clinician-researcher
- Institutional agreements are usually yours to secure, and they take longer than the modelling does.
- Imaging / data scientist
- Record the acquisition parameters, not just the pixels. They are what explains a site-to-site difference later.
- Software team
- Provenance is a data-model problem. Retrofitting it after the pipeline exists is the expensive route.
- Quality and regulatory
- This is the first thing a partner asks for and the hardest to produce late.
Who the data represents
Does the data look like the patients you intend to help?
A model can be accurate and still be inapplicable. The gap between the two populations is the scope of the claim, and it is better stated by you than discovered by a reviewer.
WHAT YOU WILL BE ASKED
- Which groups are thin or missing in your data?
- How many sites, and how different are they from each other?
- What equipment produced the images, and what will produce them in use?
- If the answer is "one site", is that a limitation or a stated scope?
START KEEPING NOW
The population description, including the exclusions you made deliberately
WHO OWNS THIS
- Clinician-researcher
- You know which populations differ clinically. That knowledge belongs in the dataset description.
- Imaging / data scientist
- Report the distribution, not only the size. A large single-site set is still one site.
- Software team
- Equipment variability becomes an integration requirement, not just a statistic.
- Quality and regulatory
- A stated scope is a defensible position; an unstated one reads as an oversight.
Which model produced that result
Could you rebuild the exact model behind the number in your deck?
Reproducibility is what turns a performance number into evidence. Without lineage, retraining silently detaches every earlier result from anything you could ship.
WHAT YOU WILL BE ASKED
- Which commit, which weights, which data snapshot?
- How many models were trained before this one, and what changed?
- If the person who trained it left tomorrow, could someone else rebuild it?
- Does the model in your deck still exist?
START KEEPING NOW
The model, its inputs and its configuration, tied to each result you rely on
WHO OWNS THIS
- Clinician-researcher
- The number you quote is only as durable as the model behind it.
- Imaging / data scientist
- This is version control for data and weights, not just code, and it is cheap now and impossible later.
- Software team
- Lineage is a build-system property. Make it automatic rather than a discipline.
- Quality and regulatory
- This is the record that makes controlled change possible at all.
How well it works, measured honestly
What has the model been tested on that it never learned from?
A set that informed any choice — a threshold, an architecture, a stopping point — is no longer independent of the result. This is the failure that survives peer review and does not survive a reviewer.
WHAT YOU WILL BE ASKED
- Which data was held back, and when was it locked?
- Has anyone looked at it more than once?
- Were the acceptance criteria written before the test or after it?
- How does performance vary by site, by scanner, by subgroup?
START KEEPING NOW
The protocol and acceptance criteria, dated before the test was run
WHO OWNS THIS
- Clinician-researcher
- Insist on criteria set in advance. It is the difference between a result and a finding.
- Imaging / data scientist
- Leakage is usually procedural rather than technical: the set was used, once, for a decision nobody logged.
- Software team
- Make the held-out set hard to touch. Access control is a better guarantee than intent.
- Quality and regulatory
- A protocol written after the result is the finding an assessor reaches for.
Whether it helps a clinician
Does a clinician using it do better than a clinician without it?
Technical accuracy and clinical benefit are different claims with different evidence. A model can be more accurate than a clinician and still not improve what the clinician does.
WHAT YOU WILL BE ASKED
- Who established the reference the model is compared against, and how?
- How many readers, how experienced, and reading under what conditions?
- Did the clinician have the model available, or the result already applied?
- What did the model change — the answer, or the time taken to reach it?
START KEEPING NOW
How the reference truth was established, and by whom
WHO OWNS THIS
- Clinician-researcher
- You design this stage. Reader-study design is where clinical credibility is won or lost.
- Imaging / data scientist
- The comparator matters more than the metric. Against what, read by whom?
- Software team
- How the result is presented is part of the intervention being tested.
- Quality and regulatory
- Keep the technical and clinical claims separate in every document.
How a person and the model work together
What does the clinician see, when, and what can they do about it?
Automation bias is a real failure mode: a reviewer defers to a proposal rather than judging it. "A human reviews the output" changes where the risk sits; it does not remove it.
WHAT YOU WILL BE ASKED
- Can the clinician see why, or only what?
- What happens when they disagree — is it easy or is it friction?
- How is uncertainty shown, if it is shown?
- Could a busy user reach a wrong conclusion by using it exactly as intended?
START KEEPING NOW
The interface as it was when each study was run — screenshots are evidence here
WHO OWNS THIS
- Clinician-researcher
- You are the only person who can say what a rushed colleague would actually do with this screen.
- Imaging / data scientist
- A calibrated confidence that nobody understands is not usable uncertainty.
- Software team
- You own this stage. The interaction is the product as much as the model is.
- Quality and regulatory
- This is human-factors risk, and it is not addressed by instructions being available.
Security and the data you hold
What would it take for someone to change what your model says, or take what it learned from?
Clinical software is subject to security expectations that scale with what it can influence, and the data behind an imaging model is among the most sensitive an organisation holds.
WHAT YOU WILL BE ASKED
- Where does patient data rest, and who can reach it?
- How would you know if a model file were replaced?
- What third-party code is in the inference path?
- What is your route to shipping a security fix once this is deployed?
START KEEPING NOW
The dependency inventory for the inference path, kept current rather than reconstructed
WHO OWNS THIS
- Clinician-researcher
- Institutional data agreements usually carry security obligations you have already signed.
- Imaging / data scientist
- Training data is an asset with an obligation attached, not just an input.
- Software team
- You own this stage, including the ability to ship a patch after deployment.
- Quality and regulatory
- Security work joins change control here; it does not sit beside it.
Where each piece of evidence goes
If someone asked for your evidence tomorrow, would you be assembling or retrieving?
Submission preparation becomes an excavation when records were never filed against the structure they would eventually need. The mapping costs little early and months late.
WHAT YOU WILL BE ASKED
- Which of your documents would you hand over unchanged?
- Which exist only as a notebook, a thread, or somebody’s recollection?
- Who would assemble it, and have they ever?
- What is missing that you already know is missing?
START KEEPING NOW
The map itself — what you hold, where it lives, and what it would support
WHO OWNS THIS
- Clinician-researcher
- Publications are not submission evidence, though they can be built from the same work.
- Imaging / data scientist
- Analyses need to be reproducible by someone else, from the record rather than from you.
- Software team
- Design and development records are read as evidence of control, not just of activity.
- Quality and regulatory
- You own this stage, and it is continuous rather than an event.
What happens when the model changes
You retrain it next year. What do you have to redo, and how would you know it got worse?
A model that changes without a controlled route detaches from the evidence that justified it. Postmarket monitoring is how a problem in the field reaches the people who can act on it.
WHAT YOU WILL BE ASKED
- What triggers a retrain, and who authorises it?
- Which evidence would you have to regenerate?
- What would you measure in the field, and who looks at it?
- How does a user tell you the output was wrong, and where does that go?
START KEEPING NOW
The change record for every model that reached a user, and what it replaced
WHO OWNS THIS
- Clinician-researcher
- Drift is often clinical rather than technical — practice changes, populations change.
- Imaging / data scientist
- Decide what you will monitor before deployment. Retrofitted monitoring measures what is easy.
- Software team
- Version what reached each user, not just what is current.
- Quality and regulatory
- You own this stage. Complaint handling and change control are the same conversation here.
A WORKED CASE
A synthetic case. It describes no real product, institution, dataset or person, and no output here is a regulatory determination about it or about anything it resembles.
A model that marks regions on chest radiographs where it estimates a pneumothorax may be present, developed from images collected during routine care at a single academic hospital.
Intended use as written: Intended to draw an emergency-department clinician’s attention to chest radiographs that may show a pneumothorax, so that those studies are looked at sooner. The clinician reads the image and makes the finding.
Their question: “We have a result we believe in and a meeting with a potential partner in four months. What do we need to have in order by then?”
STILL OPEN — RECORDED, NOT ASSUMED
- Which market is first — the answer changes the framework and nobody has written it down
- Whether performance holds at a community hospital on different equipment
- Whether the marked region is reviewable by the clinician or is effectively a black box with a box drawn on it
- Who holds the institutional agreement covering commercial use of the development images
- Still open: The prioritisation function is likely to be treated as a regulated device function in the intended market.
- Still open: Because the reader still makes the finding, the evidence expected may be lighter than for a diagnostic claim.
SIX WAYS THIS GOES WRONG
Each is one change to the case above. None requires anyone to be careless, and five of the six look like progress at the time.
The claim drifts between the deck and the protocol
The investor deck says the model "detects pneumothorax". The study protocol says it "prioritises studies for review". Nobody has noticed the two sentences differ.
Why it looks fine: Both are true descriptions of the same software, and the deck sentence is shorter and lands better in a meeting.
What it costs: Detection and prioritisation are different claims with different evidence. Promotional material is read as evidence of intended use, so the stronger sentence quietly becomes the claim the evidence has to support — and the study was designed against the weaker one.
Caught at stage 02 →The data is excellent and describes one hospital
All development images come from two scanner models at one academic centre. The team plans to deploy at community hospitals with older equipment.
Why it looks fine: The dataset is large, well-curated and clinically annotated by people who know the domain. Size feels like sufficiency.
What it costs: Accuracy on this data says little about the deployment sites. Discovered before the partnership, it is a stated scope; discovered during it, it is the finding that ends the conversation.
Caught at stage 05 →The held-out set was looked at once, for a good reason
Six months ago the team checked the held-out set to decide where to set the confidence threshold. Nothing was logged, and the set was used again for the reported result.
Why it looks fine: It was one decision, taken carefully, and the threshold is a small part of the system. Nobody trained on the data.
What it costs: The set stopped being independent at the moment it informed a choice. The reported performance is now an estimate on data that shaped the model, and the only remedy is a set nobody has touched — which usually means collecting more.
Caught at stage 07 →There is no untouched data at all
Every image has been used at some point — for training, for validation, or for the analyses that guided development. There is no set that has never informed anything.
Why it looks fine: Cross-validation has been used carefully throughout and the numbers are stable, which reads as robustness.
What it costs: Stable cross-validated numbers describe the development process, not performance on data the process has never seen. Nothing here can be presented as an independent estimate, and that is a collection problem measured in months.
Caught at stage 07 →The model is fast and arrives after the decision
In the deployment site, the clinician reads the image within two minutes of acquisition. The model returns its flag in six.
Why it looks fine: Six minutes is fast by any technical standard, and latency was never a stated requirement because nobody asked what the clinician does while waiting.
What it costs: A prioritisation output that arrives after the reading has happened influences nothing. The evidence of accuracy is intact and the product has no effect, which is the harder problem to see in a metric.
Caught at stage 09 →The model was retrained and improved
A stronger architecture was adopted and the model retrained on more recent images. The team reports better numbers and moves on.
Why it looks fine: The new model is genuinely better on every metric, and holding back an improvement feels like the wrong instinct.
What it costs: Every result generated with the previous model now describes software that is no longer the product, and no record links which model produced which number. The improvement is real; the evidence behind it has to be regenerated, and nobody planned for that.
Caught at stage 12 →THE SAME TWELVE STAGES, OTHER WORKFLOWS
Imaging is the first family SPEQ has authored, not a special case. These are the other fifteen in the same taxonomy; each changes the evidence emphasis, not the sequence. None has an authored journey yet, and this page says so rather than implying sixteen.
Digital pathology · Clinical decision support · Physiologic signal intelligence · Remote monitoring and digital measures · IVD and laboratory intelligence · Genomics and precision medicine · Surgical planning, navigation and robotics · Radiation oncology · Closed-loop therapy and device control · Patient-facing digital support · Digital therapeutics · Ambient clinical documentation and language AI · Clinical-trial AI · Pharmacovigilance AI · Health-system operations AI
SPEQ is not affiliated with, endorsed by, or acting for any regulatory authority, university or health system. This page is educational: it does not determine device status, select a regulatory pathway, state a submission type, or assess whether any model is clinically adequate. Those determinations belong to the sponsor, made with a qualified adviser against a specific framework. The national journey is at the Startup Regulatory Journey.