FDA GenAI discussion / Question 11 of 26

When can a device be confirmed without a prospective clinical study, and what earns that lighter path?

Full FDA question

CDRH is considering that clinical confirmation for a GenAI-enabled device might not require a prospective clinical study in every case, and has described above a range of approaches of increasing rigor and patient exposure. How might a sponsor select and justify a confirmation approach tailored to a device’s intended use and proportionate to the device’s risk profile? Are there device types or risk profiles for which one or more of these approaches would be insufficient or inappropriate? How might the anticipated distribution of real-world inputs be taken into consideration? Are there other methods of clinical confirmation that might help inform the evaluation of GenAI-enabled devices?
Read the FDA discussion paper ↗

26 of 95 submissions reference this question.

All audiences
16 Industry5 Clinicians1 Public / patients4 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/11
Filter by audience
Question 11 · Public feedback

What respondents recommend

13 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Check that evidence fits the intended users and setting

Proposes empirical testing of human oversight in representative workflows, without a clear prospective-study threshold.

Read the source passage
2. The risk framework: “human oversight” must be tested, not assumed Response to Discussion Questions 1, 2, and 14 The two-axis framework (device activity × severity of harm) is sound. Question 1 asks whether additional dimensions — including the time pressure of the deployment setting — should be represented. My answer is yes, and specifically: the degree of human oversight should be treated as an empirical property of the deployment setting, not as a design feature that is present or absent. In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override option does not guarantee the override will be used. Automation bias is well documented, and clinicians under time pressure defer to confident outputs. Appendix A element E.4 recognizes automation bias, but treats it as a communication-quality attribute of the device. I would encourage CDRH to also treat it as a modifier of position on the activity axis: a function nominally placed at “acts with continuous HCP supervision” may in practice operate closer to autonomy if the supervision is not exercised. I encourage FDA to: Treat “degree of human oversight” as a property to be demonstrated in representative use conditions — time-pressured, multi-patient, real interface — rather than asserted in labeling. Ask sponsors to show evidence that intended users can and do detect incorrect outputs in representative workflows. This is an override-rate and detection-rate measurement, and it is precisely the kind of human-AI team evidence contemplated in Question 14. Note that the paper’s own observation — that a “talk to your doctor” statement may not make an output less directive — applies symmetrically, and bears on Question 2: an override interface that is never used provides no oversight. Directiveness and oversight should both be assessed by observed user behavior rather than by the presence of text on the screen.
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11: Selection of the confirmation method The confirmation method should depend on the possible harm and the remaining uncertainty. A retrospective study can be sufficient for a narrow function with an objective reference result. A shadow study is useful when local data, integration, workflow, or latency can change performance. A prospective study can be necessary for an irreversible, time-critical, or high- consequence action. The manufacturer should use more than one site when population, equipment, workflow, or user behavior can affect performance. The manufacturer should report natural clinical prevalence. The manufacturer should also test enough rare and severe cases.
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Check that evidence fits the intended users and setting

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11: Selecting and justifying clinical confirmation approaches. For real-world data and data collected outside the United States, I recommend requiring a formal transportability assessment: characterization of case-mix, practice-pattern, coding, and data-capture differences between source and target settings, with quantitative adjustment (for example, standardization or reweighting) where differences are material. This is consistent with FDA's RWE program and would be a natural application of modern causal-inference and transportability methods; OUS evidence should be neither privileged nor discounted but transported transparently.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11 - Selection of clinical confirmation approach FDA's proposed ladder of approaches is appropriate. Selection should depend on intended use, risk, reversibility, degree of autonomy, availability of a valid reference standard, novelty of the deployment context, and the extent to which real-world interaction may expose behaviors not captured in static testing. Retrospective evaluation may be sufficient for lower-risk functions with stable inputs and strong reference standards. Shadow deployment is especially useful when workflow context may change system behavior but patient exposure should be avoided. Prospective studies become more important as the device independently directs or takes high-consequence action, when clinically meaningful endpoints cannot be inferred from retrospective data, or when interaction effects are central to safety.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11 — How should sponsors select clinical confirmation methods? Build an evidence ladder, and require sponsors to climb it in order. I recommend formalizing FDA's proposed progression into five tiers, with the required tier set by the risk determination in Part I: 1. Retrospective — real patient cases, no patient exposure. 2. Shadow — the AI operates in the real workflow but cannot affect care. 3. Supervised — outputs are used under defined clinician supervision. 4. Controlled autonomous deployment — the AI independently performs a narrowly defined function under enhanced monitoring. 5. Full intended-use deployment — reserved for systems whose demonstrated competence justifies the autonomy sought.
Original source ↗

Mitchell Berger

Academia / other · Aug 25, 2026

Check that evidence fits the intended users and setting

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11: Additional Clinical Confirmatory Approaches FDA Should Add: FDA’s list is strong but incomplete. Additional approaches for GenAI include: • Hallucination & Adversarial Stress Testing: Inputs designed to provoke hallucinations, misleading outputs, or unsafe reasoning. • Bias, Fairness, and Representativeness Evaluation: Testing across demographic groups, comorbidities, and underrepresented clinical conditions. • Drift and update impact analysis: Pre‑ and post‑update testing, stability testing, and drift detection. • Testing in different clinical contexts and environments: Emergency departments, inpatient, outpatient, specialty settings, and varied workflows. • Human factors testing: Provider and patient comprehension, over‑reliance, safeguard functioning, and clarity of uncertainty indicators. Sincerely, Mitchell Digitally signed by Mitchell Berger DN: cn=Mitchell Berger, c=US, email=mazruia@hotmail.com Berger Date: 2026.08.25 21:23:23 -04'00' Mitchell Berger Note: Please note that I am submitting these suggestions in my personal/private capacity. The views expressed are mine only and should not be imputed to other individuals nor to any public or private entity. 4 Page
Original source ↗

QRx Partners

Industry · Aug 24, 2026

Some uses can be confirmed without a prospective study

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11: Clinical Confirmation We support FDA's proposal that clinical confirmation be proportionate to intended use and risk and not necessarily require a prospective clinical study. The key question should be what uncertainty regarding safety and effectiveness remains after benchmarking and nonclinical evaluation, and what evidence is needed to address that uncertainty. Retrospective evaluation, shadow deployment, standardized patient interactions, clinician adjudication, and prospective investigation should be considered alternative or complementary methods rather than a fixed hierarchy. This approach supports least burdensome principles while allowing greater rigor when intended use, autonomy, potential consequences, or residual uncertainty warrant additional evidence.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11 — When should clinical confirmation include prospective testing with real patients? Response FDA should require stronger real-world confirmation as systems become more patient-facing, more autonomous, more consequential, and less subject to immediate qualified human review. Prospective testing becomes especially important when AI can:  influence emergency or urgent-care decisions;  independently interact with patients;  triage symptoms;  initiate treatment-related actions;  redirect care;  affect medication access;  influence whether a patient seeks care at all;  or carry out multiple actions before human review. There is an important difference between testing whether an AI can answer a clinical vignette correctly and testing whether it can safely interact with an actual human being who is frightened, distracted, medically unsophisticated, minimizing symptoms, worried about cost, or asking the wrong question. That difference is precisely why some patient-facing technologies should require prospective human interaction testing rather than relying entirely upon static benchmark performance.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11 — Selecting a clinical confirmation approach I support the graduated ladder. Two recommendations. First, the selection should be driven by the risk mapping recommended under Question 8, not left to sponsor justification. Sponsor discretion in selecting one’s own evidentiary bar, even with a justification requirement, produces predictable downward pressure and inconsistent review. Second, shadow deployment deserves to be the presumptive default at intermediate risk. It is the only rung that produces genuine real-world input distributions at zero patient exposure, and the divergence between benchmark input distributions and real clinical input distributions is, in my experience, where software devices most often disappoint. Retrospective evaluation on curated real inputs does not substitute, because curation removes exactly the malformed, incomplete, and atypical inputs that drive real-world failure. On input distributions: I recommend CDRH require prospective characterization of the anticipated real-world input distribution, and that confirmation evidence be labeled with the population and setting in which it was obtained. A device confirmed at three academic centers has not been confirmed for a rural critical-access 7 of 19 Docket No. FDA-2026-N-7874 hospital with different documentation practices, and the labeling should make the scope of the evidence visible to the adopting institution.
Original source ↗

Cara AI (Renee Dua, MD)

Industry · Aug 18, 2026

Some uses can be confirmed without a prospective study

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11. Selecting a clinical confirmation approach, and a note on blinded review The range of approaches described in Section V.C.1 is sound, and we support the position a prospective clinical study is not required in every case. We want to comment specifically on the blinded variant of clinician adjudication of real cases, because we are running it now. In our design, a qualified independent reviewer examines the raw inputs collected in the member's home, independently determines the correct finding, and records that determination before seeing any system output. The two are then compared. We suggest CDRH define adjudicator qualification by the task rather than defaulting to physician review. For functional assessment, home safety, and activity of daily living scoring, the appropriate expert is often a certified nurse assessor or a trained assessment coordinator rather than a physician. These are the people who hold the applicable scoring standard and who perform the task in practice. Defaulting to physician adjudication would raise the cost of clinical confirmation without improving the reference standard, and in several domains it would make the reference standard less accurate. Sequencing matters as much as credential. Once a reviewer has seen a generated output, the reviewer cannot unsee it, and agreement measured afterward is inflated by anchoring in a way invisible in the resulting statistic. We suggest CDRH make sequencing explicit in any future description of the method. A submission should state whether the reference determination was recorded before or after the reviewer saw the output, and unblinded review should not be treated as equivalent evidence to blinded review. We also note this method is achievable for a small company. It requires no research infrastructure beyond disciplined sequencing and a record of when each determination was made. If CDRH wants a form of clinical confirmation small sponsors are able to produce without dedicated trial funding, blinded paired adjudication is a strong candidate for further definition, possibly through the Medical Device Development Tool program.
Original source ↗

Richard Pescatore, DO (BellyMD)

Industry · Aug 18, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 11: proportionate clinical confirmation. I support the confirmation ladder and the recognition that a prospective study is not always necessary. For non-directive informational functions in the low-consequence region of the framework, retrospective evaluation on real-world inputs plus standardized patient interactions should ordinarily suffice, with prospective study reserved for action-directing and action-taking functions and high-consequence domains. I ask CDRH to publish presumptive confirmation tiers keyed to position on the two-axis framework, so a sponsor can locate its device and know the default evidence expectation before designing a program. The economics deserve plain statement: if every conversational function requires a prospective trial, the field consolidates to the largest incumbents, and the low-risk informational functions that would have reached unserved patients are never built. That outcome has a public-health cost too, and it is paid by patients who currently receive nothing.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 11 - Selecting clinical-confirmation rigor Trace ID. TR-Q11 | FDA Q11; Sec. V.E; App. B; pp. 18-19 / 28-29 Page 17 BCR Realization Audit - FDA GenAI Medical Devices - REV4 BCR response. Use a sequential clinical-confirmation ladder based on the highest-risk realized branch. Lower-risk informational uses may close with retrospective/shadow evidence; high-consequence or autonomous uses need stronger clinically representative confirmation and may require prospective study. BCR rule basis. BCR-R09,R10,R11,R14,R17 Solution-stack link. S7 Closure evidence. Prespecified rationale for confirmation tier; independent adjudication and representativeness evidence Pass / re-open. Selected method closes intended-use residuals at required rigor Re-open when: Risk profile, intended use, population, or unresolved residual changes.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Some uses can be confirmed without a prospective study · Require prospective studies for specified higher-risk uses

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 11: When Prospective Clinical Studies Are Required The threshold for requiring a prospective clinical study should be calibrated to actual risk, not applied uniformly across device categories. For Class II devices operating in HCP-supervised environments, shadow deployment combined with retrospective clinical outcome evaluation should generally satisfy the clinical confirmation requirement. Shadow deployment — in which the AI device's outputs are recorded alongside clinical decisions made without reliance on those outputs — provides meaningful real-world evidence without requiring the ethical and logistical complexity of a prospective randomized design. Prospective randomized clinical study design should be reserved for autonomous, patient- facing, Class III devices where AI-generated outputs will directly drive clinical actions without contemporaneous HCP oversight. Applying prospective RCT requirements broadly would render the competency-based framework economically unworkable for the vast majority of Class II generative AI device categories and would create a competitive barrier that favors only the largest manufacturers — precisely the outcome that least-burdensome principles are designed to prevent.
Original source ↗
Source directory

All 26 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Ben LocwinIndustry · Aug 26, 2026Cara AI (Renee Dua, MD)Industry · Aug 18, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026QRx PartnersIndustry · Aug 24, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Mitchell BergerAcademia / other · Aug 25, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026