FDA GenAI discussion / Question 14 of 26

For open-ended AI outputs, who is the performance comparator: a clinician panel, generalists, specialists, or the human-AI team?

Full FDA question

For open-ended device outputs, how should performance comparators and acceptance criteria be selected? When a panel of qualified clinicians serves as the comparator, how should the applicable standard (for example, the standard of care versus the performance of a median clinician in practice) be defined and justified? Should generalist or specialist physicians be used as a performance standard? When should human-AI team performance, rather than the device operating alone, serve as the basis for evaluation?
Read the FDA discussion paper ↗

25 of 95 submissions reference this question.

All audiences
15 Industry7 Clinicians1 Public / patients2 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/14
Filter by audience
Question 14 · Public feedback

What respondents recommend

12 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Navid Farr

Industry · Sep 8, 2026

Judge against the applicable standard of care · Use clinicians matched to the clinical task · Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
R7. The comparator should be the standard of care, not the "median clinician in practice" (Questions 14 and 15) This follows from S4 and S5. The paper offers two comparators: a clinician panel whose consensus reflects the standard of care, or the performance of a median clinician in practice. These are not equivalent. The median clinician in practice is, by definition, below the standard of care some of the time, and a comparator set at that level normalizes existing gaps in care as an acceptable ceiling for a new technology. A patient-first framework should set the standard of care as the floor, with prespecified non-inferiority margins justified by clinical context. Where a device is intended to extend specialist knowledge to generalists (a benefit the paper rightly names), the comparator for the device should be specialist-level performance on the specialist question, not generalist performance. Evaluation of the human-AI team should include measurement of automation bias — whether clinicians over-accept outputs — and not only measurement of the team's aggregate accuracy, since the former predicts how the team will behave when the device is wrong.
Original source ↗

Amr Saad, MD (Pallas Kliniken)

Clinicians · Sep 7, 2026

Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
This comment addresses one sentence of Discussion Question 14 and nothing else in the paper. The sentence is: "When should human-AI team performance, rather than the device operating alone, serve as the basis for evaluation?" I take the paper at its own description of its status. It "is intended for discussion purposes only and does not represent draft or final guidance," and it "is not intended to propose or implement policy changes regarding how CDRH intends to regulate generative AI-enabled devices." Nothing below is offered as a response to a proposed requirement. My answer is that human-AI team performance should be the basis for evaluation whenever a clinician who remains accountable for the resulting decision will see the device output, and that team performance cannot be measured at all unless the position of the output in the workflow is fixed as part of what is being evaluated. Section V.D.1 already approaches this. It notes that considerations for selecting the comparator "might include how the device is actually used, including the combined performance of the clinician and device working together as a human-AI team versus the device working in a fully autonomous workflow, depending on the intended use." My submission is that "how the device is actually used" is underspecified in one respect that has already produced a concrete problem in an authorized device class, and that generative outputs will reproduce it at a much larger scale. Autonomous diabetic retinopathy screening is the case in which the division of labor between device and clinician was settled by an authorization rather than by local clinical practice. In the De Novo summary for IDx- DR, CDRH states that the clinical study "demonstrated safe and effective clinical performance of IDx-DR when used to automatically (without physician assistance) detect mtmDR" (DEN180001, page 12; De Novo request received January 12, 2018, granted April 11, 2018). The pivotal trial describes the system in the same terms, as autonomous in the sense of "without human expert reading of the retinal images," and states that responsible implementation in primary care "requires autonomy (i.e., a use case that removes the requirement for review by human experts)" (Abràmoff et al., npj Digital Medicine 2018;1:39). The labeling reproduced in the same De Novo summary nevertheless assigns a task to a professional: "Physicians should review IDx-DR results and advise patients of recommended referrals to an eye care provider for evaluation and potential treatment" (DEN180001, page 2). Those two statements are compatible only if reviewing a result means something other than reviewing a finding. What reaches the physician is a classification, not the images. In the pivotal trial the system "correctly identified 173 of the 198 fully analyzable participants with fundus mtmDR," an observed sensitivity of 87.4 percent (173/198). The 25 participants in that trial who had referable disease and did not receive a positive output are known only because a reading center graded every image for study purposes. In routine use, by design, no one does. A physician who is asked to review the result and advise the patient has no means of separating that group from true negatives, so the review step named in the labeling carries a responsibility that the intended use has already made impossible to exercise. This is not a failure mode of the device. It follows from where the device was placed in the sequence, and that placement was part of what was authorized. The sequence also matters to patients, and it can be measured. Two preregistered vignette experiments with 489 and 570 members of the general public in Germany varied whether a physician used no AI support, descriptive AI support, or diagnostic AI support, and in the second study varied the timing of that support (Schaffernak et al., Journal of Medical Internet Research 2026;28:e93172). In study 2 (N=570, 7-point scales), estimated marginal means for trust in the medical decisions were 5.17 (95% CI 5.02 to 5.33) with no AI, 4.87 (4.70 to 5.04) when the physician reviewed the case and a diagnostic AI suggestion concurrently, and 5.35 (5.19 to 5.51) when the physician assessed the case independently first and reviewed the AI output afterward. Concurrent diagnostic AI was rated significantly lower than the no-AI baseline (t565=2.65; P=.04) and significantly lower than sequential diagnostic AI (t565=4.13; P<.001). The device, the output, and the clinician were held constant. Only the order changed. What I would ask CDRH to consider is the following. For a GenAI-enabled device whose output is read by a clinician who retains responsibility for the decision, the workflow position of that output should be treated as part of the device under evaluation rather than as a deployment variable, and specified in the intended use with the same precision as the output itself. Two things determine whether a human-AI team exists in the first place: t
Original source ↗

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
2. The risk framework: “human oversight” must be tested, not assumed Response to Discussion Questions 1, 2, and 14 The two-axis framework (device activity × severity of harm) is sound. Question 1 asks whether additional dimensions — including the time pressure of the deployment setting — should be represented. My answer is yes, and specifically: the degree of human oversight should be treated as an empirical property of the deployment setting, not as a design feature that is present or absent. In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override option does not guarantee the override will be used. Automation bias is well documented, and clinicians under time pressure defer to confident outputs. Appendix A element E.4 recognizes automation bias, but treats it as a communication-quality attribute of the device. I would encourage CDRH to also treat it as a modifier of position on the activity axis: a function nominally placed at “acts with continuous HCP supervision” may in practice operate closer to autonomy if the supervision is not exercised. I encourage FDA to: Treat “degree of human oversight” as a property to be demonstrated in representative use conditions — time-pressured, multi-patient, real interface — rather than asserted in labeling. Ask sponsors to show evidence that intended users can and do detect incorrect outputs in representative workflows. This is an override-rate and detection-rate measurement, and it is precisely the kind of human-AI team evidence contemplated in Question 14. Note that the paper’s own observation — that a “talk to your doctor” statement may not make an output less directive — applies symmetrically, and bears on Question 2: an override interface that is never used provides no oversight. Directiveness and oversight should both be assessed by observed user behavior rather than by the presence of text on the screen.
Original source ↗

Manuj Agarwal, MD

Clinicians · Sep 3, 2026

Judge against the applicable standard of care · Use clinicians matched to the clinical task · Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 14: Comparators and acceptance criteria The comparator should be selected according to the clinical task and intended workflow. If a device performs a function that ordinarily requires specialty expertise, a qualified specialist panel should define the reference standard. A median generalist comparator may describe current practice, but it should not define acceptable safety for a specialty task. Similarly, median clinician performance should not be used when it would normalize avoidable high-consequence errors. For many clinician-facing devices, the most informative evaluation would include four conditions: the device operating alone; the intended user operating without the device; the intended user working with the device; and an independent specialist reference panel. This design separates device capability from the real-world value and risks of the human-device team. It can also identify a device that performs well in isolation but degrades team performance by creating false reassurance, anchoring, or additional work that is difficult to verify. When human review is part of the intended use, human-device team performance should be a principal basis for evaluation, but device-alone testing should still be required to characterize latent failure modes and the burden placed on the reviewer. Oversight should not be credited as a risk control unless evaluation shows that intended users can detect representative errors within the actual time, information, and workflow constraints of use. Acceptance criteria should be prespecified and stratified by clinical consequence and case complexity. They should address not only whether a final answer is acceptable, but also whether the device omitted decisive information, recommended unnecessary or delayed care, expressed unjustified certainty, or failed to recognize that specialist input was required.
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
3.6 Question 14 - Human comparators and human-AI teams Where clinician supervision forms part of the intended use, the human-AI configuration may be the appropriate unit of clinical performance evaluation. However, “human in the loop” is not a sufficient safety claim. Oversight is meaningful only if the human has sufficient source information, sees relevant uncertainty, has a realistic opportunity to review the output, can override it and is not pushed into automation bias by interface design or workflow pressure. This position aligns with the international GMLP principle that attention should be paid to human-AI team performance and the transparency principle that users need clear, essential information to make informed decisions (FDA, Health Canada and MHRA, 2024; IMDRF, 2025).
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Judge against the applicable standard of care · Use clinicians matched to the clinical task · Evaluate the clinician and AI working together

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 14 - Comparators and acceptance criteria for open-ended outputs The comparator should match the device's intended clinical role. A device intended to assist a generalist should not automatically be judged only against a subspecialist, and a device 8 intended for autonomous specialty decision-making should not be validated merely by comparison with an unaided generalist. The relevant comparator may be a qualified clinician panel, a validated reference standard, a human-AI team, or a combination depending on intended use. Acceptance criteria should be multidimensional where a single "correct" response does not exist. Relevant dimensions can include safety-critical omissions, factual correctness, appropriateness of differential reasoning, calibration, escalation behavior, communication quality, and consistency across equivalent scenarios.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Judge against the applicable standard of care

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 14 — Against what should AI performance be compared? Response FDA should be cautious about treating the performance of the “average” clinician as the ultimate standard against which AI is judged. Average practice and evidence-based best practice are not necessarily the same thing. Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 10 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices Clinical variation exists for many reasons: differences in training, experience, local custom, available resources, cognitive bias, information overload, financial incentives, and simple lag between new evidence and widespread adoption. Therefore: AI should not be considered safe merely because it performs as well as the average clinician if the average clinician is not consistently following the best current evidence. The preferred comparator should generally include a current, evidence-based clinical standard. This is an area where AI has extraordinary potential. No individual physician, nurse practitioner, physician assistant, pharmacist, nurse, or other professional can personally remain current on every relevant study, guideline, therapeutic development, contraindication, drug interaction, diagnostic refinement, and care pathway across modern medicine. The volume is beyond human capacity. That does not make clinicians obsolete. It makes AI potentially indispensable. Good AI can help a practitioner retrieve information. Great AI can continuously place current evidence beside human judgment and make unexplained departure from that evidence visible. That may protect clinicians from:  anchoring;  availability bias;  outdated training;  overconfidence;  local custom;  premature closure;  excessive deference to personal experience;  and the natural human tendency to believe that expertise accumulated over many years remains sufficient by itself. There is an uncomfortable but important point here. Professional expertise can sometimes become self-reinforcing. A highly experienced clinician may believe, “I know how to treat this,” when the evidence has changed since the practice pattern was learned. AI can help keep that human tendency in check. It can effectively ask: “Before we proceed, is this still what the evidence says?” That is one of AI’s greatest potential contributions to healthcare. FDA should therefore consider evaluation models that compare AI-supported care against: 5. the best available evidence-based standard; 6. actual current unaided practice; and 7. the performance of the human-AI team. Those three comparisons answer different and important questions.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Judge against the applicable standard of care

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 14 — Comparators and acceptance criteria I recommend against “the median clinician in practice” as a primary performance standard, for several reasons. It is unmeasured — there is no reliable characterization of median clinician performance for most tasks a GenAI device would perform, so the comparator would in practice be a convenience sample presented as a population parameter. It is a moving target, and one that a widely adopted device would itself move. Most importantly, it anchors the acceptable bar to observed practice variation rather than to patient benefit: where median practice is poor, the standard licenses a device that is also poor, and the patient in front of that device receives no protection from the fact that the alternative was equally bad. The better construction, in my view, has two parts: 1. Prespecified clinical acceptance criteria derived from the consequence of the decision. What performance is required for this device, in this use, to be safe and effective, given what happens when it is wrong? This is a clinical and normative judgment, it should be made in advance and defended, and it is the judgment the acceptance decision actually turns on. 2. A clinician panel as an adjudication instrument, not as the bar. Panels are essential for open-ended outputs, where automated scoring is inadequate. But a panel is a measurement instrument, and it should be validated and reported as one: prespecified composition and qualifications, prespecified size with 9 of 19 Docket No. FDA-2026-N-7874 justification, a written adjudication protocol, blinding to output source where feasible, measured and reported inter-rater reliability, and a prespecified procedure for resolving disagreement. A panel with unreported inter-rater reliability is an uncharacterized instrument, and evidence generated by an uncharacterized instrument should not gate a marketing decision. This is standard practice in reader studies and should carry over without modification. On generalist versus specialist: the comparator should match the labeled intended user, subject to my comment under Question 4 regarding foreseeable use. A device labeled for primary care benchmarked against subspecialists is being measured against a standard its users do not represent — which may overstate or understate the device’s contribution depending on direction, and in either case measures the wrong thing. On human-AI team performance: where the labeling requires human review, team performance should be the primary basis of evaluation, and standalone device performance should be supporting evidence only. Evaluating the device in isolation when it will never be used in isolation measures a configuration that does not exist. I would add that team evaluation must be designed to detect automation bias — the tendency of reviewers to defer to a confident output — which is a principal mechanism by which human-in-the-loop safeguards fail. A team evaluation that does not include cases where the device is confidently wrong will systematically overstate the safeguard’s value.
Original source ↗

Cara AI (Renee Dua, MD)

Industry · Aug 18, 2026

Compare with what happens without the device

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 14. Comparators for open-ended outputs Comparing a system against a panel of qualified clinicians assumes a stable human reference standard exists. In Medicaid functional assessment, it largely does not. © 2026 Cara AI, Inc. The instrument itself is not standardized. The federal government does not require states to use a particular functional assessment tool, and a national inventory identified at least 124 tools currently in use.1 The variation is not cosmetic. The District of Columbia bathing assessment collects the frequency and duration of assistance required, while the Kentucky assessment does not.2 Individual health plans then apply their own scoring conventions and thresholds within their state's requirements. A reference panel drawn from more than one program may disagree not because clinical judgment differs but because the reviewers are applying different instruments. Agreement among assessors under real field conditions is also largely unmeasured. Published inter- rater reliability studies for activity of daily living instruments are generally conducted in controlled research settings, with trained raters applying a single instrument to a small sample, and several report high agreement. Those results do not establish that assessors agree during Medicaid home visits, across different instruments, under time pressure, in the member's own environment. We are not aware of a body of evidence establishing a reliable human reference standard for this task in the deployment setting, and we suggest CDRH treat the existence of that standard as something a sponsor demonstrates rather than assumes. Two suggestions follow. Where the reference standard is expert judgment rather than an objective result such as a biopsy, sponsors should report inter-rater agreement within the human reference panel alongside system-to-panel agreement, and should state which instrument each reviewer applied. A system matching a panel agreeing with itself 70 percent of the time is a different finding from a system matching a panel agreeing with itself 95 percent of the time, and a raw agreement figure hides the difference. Where the purpose of a function is to reduce variance rather than to outperform a median clinician, consistency should be measurable as a performance endpoint in its own right. In authorization- driven programs, two assessors reaching the same score for the same member is itself the clinical and equity objective.
Original source ↗

Richard Pescatore, DO (BellyMD)

Industry · Aug 18, 2026

Compare with what happens without the device

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Questions 14 and 15: comparators. This is the most consequential question in the paper. For informational functions deployed directly to patients, a specialist panel is often the wrong comparator because it does not describe what the device displaces. In much of DGBI care, the device does not replace a specialist; it replaces nothing, or it replaces uncurated internet content. I recommend: first, comparator selection anchored to the realistic deployment context, supported by care-access data for the intended population; second, recognition of a "usual information environment" comparator, validated by sampling what patients in the intended population encounter when they search their symptoms; third, reservation of clinician-panel parity for functions that displace clinician judgment, meaning action-directing and action-taking functions or deployment inside clinical workflows; and fourth, where clinician comparison is appropriate for generalist-shaped contexts, the median generalist rather than a specialist panel. Holding low-risk educational software to a standard the delivery system itself does not meet protects no one. It preserves the status quo for patients whose status quo is nothing.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Judge against the applicable standard of care · Evaluate the clinician and AI working together · Compare with what happens without the device

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 14 - Comparator and acceptance criteria Trace ID. TR-Q14 | FDA Q14; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. Choose the comparator from intended use and counterfactual workflow. Use an appropriate clinical standard for substitution claims; test human-AI team and device-alone performance separately when both matter. BCR rule basis. BCR-R01,R08,R10,R17 Solution-stack link. S1,S7 Closure evidence. Comparator rationale tied to intended use/counterfactual; device-alone and human-AI results separated when relevant Pass / re-open. Comparator directly supports claimed role and clinical question Re-open when: Intended-use/workflow/comparator standard changes.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Judge against the applicable standard of care

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 14: Performance Comparators — Against the Right Standard The choice of clinical performance comparator is consequential and must be handled carefully. WHM recommends against using "median clinician in practice" as the primary performance comparator for specialty AI devices. This standard, while operationally measurable, normalizes the level of care that happens to prevail in clinical practice — which may reflect training gaps, resource constraints, or systemic underperformance rather than optimal care. Clearing an AI device because it performs at the level of the median clinician in practice does not ensure patient benefit; it may simply replicate existing suboptimal performance at scale. For specialty devices, the appropriate comparator is board-certified specialist performance within the device's intended clinical domain and under conditions that reflect intended-use clinical environments. For devices intended to expand access to specialist-level care in underserved settings, comparison against delayed specialist review — reflecting the realistic clinical alternative — is appropriate and clinically meaningful. FDA should codify these comparator options in its decision matrix and guidance documents.
Original source ↗
Source directory

All 25 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Cara AI (Renee Dua, MD)Industry · Aug 18, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Amr Saad, MD (Pallas Kliniken)Clinicians · Sep 7, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Martin HaimerlAcademia / other · Sep 1, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026