FDA GenAI discussion / Question 7 of 26

Is the two-step approach, benchmark the AI, then confirm it in clinical use, the right way to evaluate these devices?

Full FDA question

Is the competency-based approach described above, i.e., device benchmarking followed by clinical confirmation, a useful and appropriate framework for evaluating GenAI-enabled devices?
Read the FDA discussion paper ↗

23 of 95 submissions reference this question.

All audiences
15 Industry3 Clinicians2 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/7
Filter by audience
Question 7 · Public feedback

Positions on this question

8 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13.

Use that sequence, with changes or conditions7
Use benchmarking followed by clinical confirmation1
Walnut Hill MedicalIndustry · Aug 18, 2026
Do not use that evaluation framework0

No analyzed submissions in this group.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Navid Farr

Industry · Sep 8, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
R5. The clinician-credentialing analogy justifies competency testing; it does not justify relaxed change control or diffused accountability (Question 7) This also follows from S3, and it is the point on which I most want to be precise, because the analogy is doing a great deal of work in Section V and its limits are not stated. The competency-based approach is described as "inspired, at a high level, by how human clinicians are evaluated and credentialed." As a justification for structured assessment of knowledge, reasoning, and safety behavior, followed by supervised practice, the analogy is apt and I support the resulting framework. But the analogy breaks in four places, and each break has a regulatory consequence. Identity and continuity. A credentialed clinician is one person whose competence, once assessed, persists until the next assessment. A "credentialed" GenAI device is a configuration that can be altered — by the sponsor or by an upstream developer — at any time, in every deployed instance simultaneously, without anyone re-examining it. A clinician cannot be silently replaced by a different clinician between board certification and practice. A model can. The regulatory consequence is that competency evidence must be bound to a specific, immutable configuration, and any change to that configuration must trigger reassessment (see R4). The analogy supports competency testing; it does not support the flexibility on postmarket modification that the paper explores in Section VI.C. Accountability. The paper cites Freyer et al. for the observation that clinicians "face professional, legal, and reputational consequences when expected standards are not met" and that "analogous mechanisms are needed for flawed GenAI-enabled devices." It then does not propose any. A clinician who harms a patient can lose a license. A device that harms a patient should face a functionally equivalent mechanism: a defined process for suspending authorization, with prespecified triggers tied to the postmarket thresholds in R2, that is faster than a recall and does not depend on the sponsor's willingness to act. Without it, the credentialing analogy grants the privileges of licensure without its obligations. Docket No. FDA-2026-N-7874 — Individual comment — Page 5 Scale. A clinician's error affects one patient at a time and is bounded by how many patients that clinician sees. A systematic error in a deployed model affects every patient who presents with the triggering pattern, everywhere, at once, until detected. Individual credentialing evolved for individual-scale risk. The regulatory consequence is that acceptance criteria for safety-critical behaviors (S.1, S.2, S.3) should be stricter for a device than would be tolerated in a single practitioner, because the harm from a given failure rate is multiplied by deployment scale. Judgment about one's own limits. Credentialing assumes the credentialed party can recognize when a case exceeds their competence and refer it onward. The paper's S.3 (calibration and clinical deferral) tests for this, which is correct. But a clinician's uncertainty is grounded in an understanding of the case; a model's expressed uncertainty is a generated output that can be as confabulated as any other. Deferral behavior should therefore be evaluated adversarially and across paraphrases (as R.1 contemplates), and should not be accepted as a mitigating safeguard on the strength of the device stating that it is uncertain. None of this argues against the competency-based framework. It argues that the framework should be adopted for what the analogy supports — structured, proportionate assessment of a fixed configuration — and that its use to justify lighter change control or shared accountability should be rejected explicitly, so that the analogy is not stretched later in guidance or in individual submissions.
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
3.4 Question 7 - Competency-based premarket evaluation The FDA explores a competency-based approach that combines non-clinical device benchmarking with clinical confirmation (FDA, 2026a). Sitora supports this direction because exhaustive testing of every possible open- ended interaction is unrealistic. However, competency must attach to the complete deployed system, not merely the foundation model. The same underlying model can perform differently when combined with different prompts, retrieval architectures, tool permissions, clinical rules, interfaces, memories and escalation policies. Strong performance by a foundation model on a medical benchmark therefore cannot, by itself, demonstrate the safety of a downstream medical device. The regulatory question should be whether the configured system can safely and effectively perform its claimed clinical function in representative conditions. This is consistent with total product lifecycle principles and clinically relevant testing in GMLP (FDA, 2025a; IMDRF, 2025).
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 7 - Usefulness of competency-based benchmarking plus clinical confirmation The proposed competency-based approach is useful and appropriate as a high-level framework, and is consistent with the broader regulatory discussion around clinician- inspired competency, lifecycle assessment, and accountability [8-11], particularly because FDA proposes evaluating the final user-facing device as configured and intended to be deployed rather than the foundation model standing alone [1]. That focus is important: clinically relevant behavior emerges from the combination of the application, model, prompts, retrieval, tools, context, interface, and workflow. Competency evaluation should nevertheless be treated as necessary but not sufficient. It establishes whether the configured device can perform the intended task under defined conditions. It does not by itself establish that the correct patient context was used, that the evidence was current, attributable, and appropriate for the contemplated clinical use, that the postmarket configuration remained equivalent to the tested configuration, or that a downstream action was authorized. The assurance functions described in this comment should not be understood as a uniform regulatory burden. Their relevance should be proportional to intended function, degree of autonomy, clinical consequence, external dependencies, and execution authority. A bounded advisory function may require only a subset of these controls; a consequential 6 autonomous function may require substantially more of them to establish an appropriate level of assurance.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 7 — Is benchmarking followed by clinical confirmation the right sequence? This is the right idea. Extend it past the day of clearance. Benchmarking → Clinical Confirmation is directionally correct, but it should be the front end of a longer pipeline. I recommend the same nine stages detailed in the Integrating Framework section below: Development → Competency Examination → Clinical Confirmation → Supervised Clinical Deployment → Provisional Authorization → Independent Authorization → Continuous Monitoring → Revalidation → Expansion or Restriction. No model should need to demonstrate every possible clinical scenario before deployment — and no model should be assumed permanently competent because it passed one benchmark. This is the central recommendation of this comment.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 7 — Is competency benchmarking followed by clinical confirmation an appropriate framework? Response Yes, provided competency is defined broadly enough. Healthcare competency cannot mean merely possessing medical knowledge or producing factually correct answers. A patient-facing AI should demonstrate competency in at least four different domains: 1. Knowledge — Does it know the relevant clinical evidence? 2. Reasoning — Can it apply that evidence appropriately to the individual situation? 3. Recognition — Can it identify danger, uncertainty, missing information, and situations requiring escalation? 4. Restraint — Does it know when it should not act independently? Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 8 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices The fourth category may ultimately be among the most important. A highly capable agentic system can potentially cause more harm, not less, if increased capability is paired with misplaced confidence or inappropriate autonomy. FDA’s proposed competency framework appropriately includes issues such as safety-critical recognition, escalation, uncertainty, ambiguous presentation, communication, and scope boundaries. I strongly encourage FDA to add a specific competency for: recognizing latent clinical need within apparently nonclinical or administrative interactions. This would directly address the type of failure identified in PatientAgentBench and increasingly likely to arise as patient-facing agents gain access to scheduling, benefits, pharmacy, communications, and other healthcare tools.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 7 — Is the competency-based approach appropriate? Broadly, yes, and the paper’s framing is a genuine contribution. But the analogy carries an assumption that should be made explicit before it is relied upon, because I believe the analogy is doing more argumentative work than it can support. Physician credentialing is evidentiarily permissive — no exhaustive testing of every scenario — because it is embedded in a system of continuous individual accountability: ongoing licensure with revocation authority, institutional privileging and peer review, malpractice liability, professional norms, and a career-long incentive to notice and correct one’s own errors. The credentialing exam is not the safety mechanism. It is the entry gate to a system whose safety mechanisms operate continuously thereafter. A device inherits none of this. There is no analogue to license revocation for a deployed model, no peer review of its individual decisions, no professional identity that responds to being wrong. The manufacturer’s quality system is the closest analogue, and it is a genuine one, but it is a different mechanism with different failure modes. There is a second disanalogy worth stating. Physicians generalize from training in ways whose failure modes are broadly legible to other physicians — we have vocabulary for the kinds of errors a tired resident makes. Generative model failures are not legible in the same way. They can be sharp, input-specific, and invisible to any amount of aggregate performance measurement. My recommendation is that CDRH retain the competency framing as an organizing structure — it maps cleanly onto benchmarking, confirmation, and monitoring — while explicitly declining to import its evidentiary permissiveness. The paper should state that the licensure analogy justifies the shape of the evidence, not the quantity, and that the reduced-testing feature of physician credentialing is a consequence of accountability structures the device context does not reproduce.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Use that sequence, with changes or conditions

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
FDA Question 7 - Competency-based premarket approach Trace ID. TR-Q07 | FDA Q7; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. The competency-based model is useful if benchmark -> clinical confirmation -> postmarket monitoring forms one continuous closure chain tied to the exact final deployed configuration. BCR rule basis. BCR-R01,R03,R09,R15,R17 Solution-stack link. S0,S5,S6,S7,S12,S13 Closure evidence. Traceable benchmark + clinical confirmation + residual monitoring for exact deployed configuration Pass / re-open. Evidence chain remains valid to deployed configuration and essential residuals close Re-open when: Any safety- relevant configuration or intended-use change.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Use benchmarking followed by clinical confirmation

Classified on this question’s stated measure only; no position is imported from other questions.

Read the source passage
Response to Question 7: The Competency-Based Approach Is Appropriate Yes. The competency-based approach is the appropriate framework for evaluating generative AI devices, and FDA deserves credit for developing it. It mirrors the structure by which human clinical expertise is validated — board certification, specialty credentialing, simulation-based assessment — and creates a principled basis for evaluating AI systems that operate similarly to clinical judgment rather than to traditional deterministic software. WHM strongly supports this framework as the foundation for FDA's regulatory approach.
Original source ↗
Source directory

All 23 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026AnonymousIndustry · Aug 18, 2026Ben LocwinIndustry · Aug 26, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Shannon KamalakerClinicians · Aug 19, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026