FDA GenAI discussion / Question 15 of 26

Could the AI be measured against what would have happened without it: unaided judgment, a delayed specialist, or no intervention?

Full FDA question

Are there ways in which performance might be assessed relative to the care, technology, or course of action likely to occur in the absence of the device, rather than to the comparators described in this section? What approaches might be used to identify and justify a comparator such as unaided clinical judgment, delayed specialist review, or no intervention?
Read the FDA discussion paper ↗

23 of 95 submissions reference this question.

All audiences
16 Industry4 Clinicians1 Public / patients2 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/15
Filter by audience
Question 15 · Public feedback

Positions on this question

9 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Sam Rosenthal (Red Kit)

Industry · Sep 9, 2026

Use that comparator, with conditions

Recommends the untrained-bystander baseline first, while retaining other comparisons; the setting is emergency protocol delivery.

Read the source passage
Question 15 — the comparator in the absence of the device. For my setting the honest comparator is not a clinician panel and not a median clinician; it is an untrained person with no signal, acting from memory. Bystander CPR rates, correct-technique rates among lay rescuers, and the outcome of "did nothing" are all published, and a device should be judged first against that baseline and only second against the standard of care it is trying to deliver. A framework that measures a bystander tool only against a clinician will conclude it is unsafe; a framework that measures it against the alternative the user actually has will ask the right question, which is whether it moves an untrained person closer to the published protocol than they would get on their own. I would ask CDRH to state that for patient-facing functions whose intended use is explicitly "when no professional is available," the primary comparator is the unaided user.
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Use the likely care without the device as a comparator

Accepts the no-device comparator for assistive communication, where no device means no equivalent expressive function. Explicitly takes no position on access-limited clinical alternatives.

Read the source passage
Summary of position Yes, and for one class of device it is the only comparator that describes reality. For augmentative and alternative communication, the absence of the device is not a slower or less expert version of the same function. It is the absence of the function. There is no unaided clinical judgment to fall back on, because the judgment being aided is the user's own expression, and no clinician can supply it. We submit one structural point that we believe is not yet reflected in Section V.C, and that the absence-of-device comparator makes visible where the panel and standard-of-care comparators do not: • Device performance is not bounded below by the no-device baseline. A comparator framework built on panels or standard of care implicitly treats the device as adding some amount of benefit between zero and the expert ceiling. A generative device that fabricates output does not land at zero. It lands below it, because the no-device condition produces no false record and the device does. • In field measurement on 2026-09-03, 39 of 136 transcribed words — twenty-nine percent — were produced from audio windows carrying no live signal. In the absence of the device those words do not exist. With the device they exist, are fluent, and are attributed to the speaker. • For this population the fabricated sentence is not a degraded output. It is an utterance entered into the record under the user's name, by a user who cannot contest it. 1. What “absence of the device” means when the function is expression Section V.C offers unaided clinical judgment, delayed specialist review, and no intervention as candidate comparators. The first two assume a human professional performs the function more slowly or less expertly without the device. That assumption holds for a diagnostic aid. It does not hold here. For a nonverbal user, the function is speech. In the absence of the device there is no delayed specialist, no unaided equivalent, and no slower path to the same output. The comparator is the third one on CDRH's list — no intervention — and for this population no intervention has a specific operational meaning: the user's FDA-2026-N-7874 — Question 15 1 The Christman AI Project intent is not expressed at all, and no record of it enters the world. This has a consequence for how benefit is measured. Against a panel comparator, a device that produces a plausible but wrong output scores as a partial success. Against the absence comparator, the same output is a categorical harm, because the counterfactual is not a worse sentence. It is no sentence. 2. The measured case, and why it inverts the sign The evidence below is drawn from the same measurement record submitted with our comment on Question 24 and is available to CDRH on request. Three screen recordings were captured between 01:27 and 02:42 on 2026-09-03 while dictating into the built-in speech-to-text of a commercial AI assistant application. Audio was extracted to 16 kHz mono PCM and examined in contiguous 250 ms windows across the full duration of each file, measured by the count of distinct 16-bit sample values per window. Live speech windows in these files carry 7,000 to 12,000 distinct values. Windows in which capture had failed collapse to single digits, with one constant held across thousands of consecutive samples. Loss rose across the session: Recording Audio carrying no signal Share of file REC1 14.00 s of 197.4 s 7.1% REC2 40.25 s of 210.4 s 19.1% REC3 31.00 s of 73.5 s 42.2% Transcribed with a widely deployed open-weights speech recognition model under the submitter's control, REC3 produced 136 words. Thirty-nine of them fall inside windows carrying no live signal. One sequence, spanning 20.00 to 25.25 seconds — six distinct sample values, 99% of samples on a single constant — was rendered as the phrase “stopped me right,” which reassembles into the speaker's own description of the failure. In a second recording the transcriber produced a grammatical sixteen-word sentence across a dead region the speaker had not spoken into. Now apply each comparator to that sixteen-word sentence. Comparator How the fabricated sentence scores Panel of qualified clinicians A wrong output. Scored against a correct one; partial credit possible. Standard of care Below standard. Still on the scale. Median clinician in practice An error of the kind a human might also make. Comparable. Absence of the device An utterance that would not exist. No partial credit is available; the counterfactual contains nothing to be worse than. Only the fourth row records what actually happened. The first three describe a device that underperformed. The fourth describes a device that manufactured a fact. 3. Why the absence comparator cannot simply replace the others We are not proposing that absence of the device become the primary comparator for GenAI-enabled devices generally. It is a poor instrument for measuring benefit. Against no device at all, almost any functioning FDA-2026-N-7874 — Question 15 2 The Christman AI Project system looks excellent, and a sponsor permitted to choose it would be choosing the weakest available bar. Its value is the opposite of a benchmark. It is a floor test. The panel and standard-of-care comparators answer how good is this device. The absence comparator answers a different and narrower question: does this device ever produce an output that is worse than nothing. Those are not the same measurement and a device can pass the first while failing the second. We therefore recommend the absence comparator be used alongside, not instead of, the comparators in Section V.C — and specifically as a pass/fail gate rather than a scored dimension. A device that produces fabricated output over null input has failed regardless of how well it performs on the panel comparator, because the two results are measuring different things and the good one does not offset the bad one. 4. Identifying and justifying the comparator Responsive to the second half of Quest
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Discussion Questions 14–15 – Reference standards and clinically relevant comparators Regarding performance comparators, I suggest more clearly separating the role of the reference standard from the role of the clinical comparator. Where a clinical comparator is used, it should reflect a relevant alternative to the device rather than being conflated with the reference standard. According to standard rules for medical devices, the GenAI device needs to be compared to the performance of the established standard-of-care to gain market access. This means, that it needs to be assessed how the GenAI system performs in relation to the standard-of-care when comparing both outcomes to an idealized reference standard or ground truth. The reference standard itself should establish, as reliably and accurately as possible, what diagnosis, assessment, or course of action is correct or clinically appropriate for the particular test case. Instead, the comparator, i.e. standard-of-care, should represent the care that would realistically occur in the absence of the device, including the intended user group and clinical environment. The Discussion Paper already raises this possibility in Question 15, including unaided clinical judgment, delayed specialist review, or no intervention. Accordingly, the experts establishing the reference standard need not be the clinicians against whom the device is compared. For a device intended to support generalist physicians, for example, specialist adjudication may establish the reference standard while representative generalist physicians provide the clinically relevant comparator. Similarly, for patient-facing home-use devices, representative patients or lay users may be the relevant user comparator, while the reference standard may still rely on specialist adjudication or another high-quality clinical reference. This approach shifts the focus from an absolute comparison of GenAI with an expert, who may not represent the intended user, toward a relative assessment in which the GenAI-enabled device and the intended user group are compared by reference to the same reference standard. Based on this, the demonstration of non-inferiority to the comparator as a relative criterion gets the main objective instead of determining an absolute level of deviation between the GenAI and the reference standard. The evaluation should also generally focus on the combined human-AI system rather than comparing the stand- alone GenAI-enabled device with the intended user. More generally, the following clinical comparison may be appropriate: • intended user + standard of care + GenAI device versus • intended user + standard of care without the GenAI device, with both evaluated against the same independent reference standard. This provides a more direct assessment of whether the device improves or at least preserves decision quality within its intended context of use than a simple comparison of stand-alone GenAI output with expert opinion. Relative non-inferiority or superiority approaches may help avoid arbitrary absolute performance thresholds, although critical safety outcomes should still be subject to absolute risk-based acceptance criteria. This approach also creates a clearer distinction between the information used to establish the reference standard and the information available to the evaluated users or device. While subsequent diagnostic findings or follow-up may strengthen the reference standard, the investigational and comparator conditions should only have access to information that would realistically have been available at the relevant decision point. Overall conclusion for Section V In summary, I support the proposed combination of competency-based benchmarking and clinical confirmation, but suggest strengthening the framework, in particular, in the following areas. • better discriminate between evaluating errors in the output of the GenAI device and consequences that result from these errors; • more explicitly link evidence intensity to criticality while using Product-Specific Risk Management to determine the specific content of evaluation; • expand benchmarking from predominantly dataset-based assessment toward dynamic, risk-based evaluation suites; • establish the use of an independent reference standard where the GenAI device can be compared to the established standard of care in a relative way; • distinguish intrinsic device performance from performance of the device within the intended human and clinical system; • treat evaluation as a lifecycle process in which critical assumptions, safeguards, and the authorized operating envelope can be progressively confirmed, expanded, and periodically reassessed. Such an approach would preserve the scalability sought by a competency-based framework while providing a clearer link between device competence, risk-control effectiveness, clinical decision quality, and ultimately reasonable assurance of safety and effectiveness. Section VI – Postmarket Mon
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 15 - Comparison with likely care in the absence of the device Yes. In some settings the clinically relevant comparator is the counterfactual care pathway rather than an idealized expert. A device may create benefit by accelerating access to specialist-level review, improving triage consistency, reducing delay, or increasing detection even if it does not outperform the best available specialist on every case. Sponsors should justify the comparator based on the intended deployment context and should avoid choosing a weak comparator merely to demonstrate superiority. Where practicable, analyses should distinguish device-alone performance, human-alone performance, and the performance of the intended human-AI workflow.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 15 — Should AI sometimes be compared with what would occur without it? Response Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 11 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices Yes. FDA should understand both absolute performance and incremental improvement. Healthcare should not reject a beneficial AI system merely because it is imperfect when the realistic alternative is slower, less consistent, less evidence-based, or unavailable. At the same time, a system should not be treated as adequately safe merely because existing care is poor. A useful comparison would therefore examine: 8. Evidence-based ideal — What should happen according to the best current clinical evidence? 9. Real-world baseline — What actually happens today without the AI? 10. AI-assisted pathway — What happens when the technology is introduced? This framework permits FDA to distinguish between two very different claims: “This AI is better than what frequently happens today.” and “This AI meets an acceptable clinical safety standard.” Both matter. They are not the same.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 15 — Counterfactual comparators I support permitting comparison to the care that would occur absent the device, and I think the paper is right that this is often the clinically meaningful question. For access-expanding devices, the true alternative is frequently no assessment at all, delayed assessment, or assessment by a less qualified party, and holding such a device to a specialist standard can foreclose a real benefit. Two safeguards should accompany it. First, the counterfactual must be empirically characterized for the specific deployment setting, not assumed — the claim that “nothing would otherwise happen” is an evidentiary claim about a care pathway and should be supported. Second, and more importantly, the comparator must travel with the labeling. A device authorized against a no-intervention counterfactual in a resource-limited setting has not been shown safe and effective as a substitute for specialist review in a resource-rich one, and the authorization should say so in terms an adopting institution can act on. Without this, counterfactual comparators become a low-bar entry route followed by scope creep that no one has evaluated.
Original source ↗

Cara AI (Renee Dua, MD)

Industry · Aug 18, 2026

Use the likely care without the device as a comparator

Classified on this question’s stated measure only; no position is imported from other questions.

Read the source passage
Question 15. Comparison against the alternative existing today We support this framing strongly and believe it deserves more weight than the paper currently gives it. In the home assessment setting, the realistic alternative to a generative AI-supported assessment is not a careful expert evaluation. It is a paper form completed under time pressure, sometimes partly from recall after the visit ended, frequently incomplete, and drawn from an instrument varying by state and by plan. The prevailing alternative is not only lower quality than an expert evaluation. It is also unstandardized, which means it produces different results for similar members. Federal improper payment reporting attributes the large majority of Medicaid improper payments to 1 1 Medicaid and CHIP Payment and Access Commission, “Functional Assessments for Long-Term Services and Supports,” Chapter 4, Report to Congress on Medicaid and CHIP, June 2016. 2 Ibid., Box 4-2. © 2026 Cara AI, Inc. insufficient documentation and missing administrative steps rather than to fraud.3 That is the operative baseline in this setting. The downstream effect is visible in service delivery. Because no reliable shared record exists of what a member was already assessed for and already received, duplication and gaps occur side by side. One member accumulates several pieces of the same durable medical equipment while another with the same documented need receives none. This is a consequence of unstandardized and incomplete assessment rather than of clinical disagreement, and it is the kind of outcome an evaluation anchored to an idealized standard of care will not surface. Evaluating a system against an idealized standard of care while the real alternative is an incomplete form produces a systematically misleading benefit-risk assessment, and it does so in a direction disadvantaging the populations with the least access to care. We suggest CDRH permit sponsors to characterize the actual prevailing practice in the deployment setting, supported by evidence, and to evaluate against it in addition to any clinical reference standard.
Original source ↗

Richard Pescatore, DO (BellyMD)

Industry · Aug 18, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Questions 14 and 15: comparators. This is the most consequential question in the paper. For informational functions deployed directly to patients, a specialist panel is often the wrong comparator because it does not describe what the device displaces. In much of DGBI care, the device does not replace a specialist; it replaces nothing, or it replaces uncurated internet content. I recommend: first, comparator selection anchored to the realistic deployment context, supported by care-access data for the intended population; second, recognition of a "usual information environment" comparator, validated by sampling what patients in the intended population encounter when they search their symptoms; third, reservation of clinician-panel parity for functions that displace clinician judgment, meaning action-directing and action-taking functions or deployment inside clinical workflows; and fourth, where clinician comparison is appropriate for generalist-shaped contexts, the median generalist rather than a specialist panel. Holding low-risk educational software to a standard the delivery system itself does not meet protects no one. It preserves the status quo for patients whose status quo is nothing.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Use that comparator, with conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
FDA Question 15 - Compare to care absent the device Trace ID. TR-Q15 | FDA Q15; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. Yes - include the actual counterfactual care likely without the device: unaided clinician, delayed specialist review, alternative technology, or no intervention when clinically appropriate. BCR rule basis. BCR-R01,R09,R10,R17 Solution-stack link. S1,S7 Closure evidence. Prespecified credible counterfactual: unaided clinician, delayed specialist, alternative technology, or no intervention as appropriate Pass / re-open. Counterfactual aligns with claimed benefit-risk statement Re-open when: Care pathway or standard practice changes.
Original source ↗
Source directory

All 23 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Ben LocwinIndustry · Aug 26, 2026Cara AI (Renee Dua, MD)Industry · Aug 18, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Joel GrunhutPublic / patients · Sep 7, 2026Martin HaimerlAcademia / other · Sep 1, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026