FDA GenAI discussion / Question 9 of 26

Do the ten benchmark competencies, from clinical knowledge to generalizability, add up to enough evidence of safety and effectiveness?

Full FDA question

Would a benchmarking structure such as the one described above be likely to provide adequate evidence of clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability to support a reasonable assurance of safety and effectiveness? Are there elements that are missing, redundant, or inappropriately categorized? Are there externally developed standards that could be leveraged?
Read the FDA discussion paper ↗

31 of 95 submissions reference this question.

All audiences
Alfred McBrideIndustry · Aug 18, 2026Ben LocwinIndustry · Aug 26, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026The Christman AI ProjectIndustry · Sep 8, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Martin HaimerlAcademia / other · Sep 1, 2026Mitchell BergerAcademia / other · Aug 25, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026
20 Industry6 Clinicians2 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/9
Filter by audience
Question 9 · Public feedback

Positions on this question

14 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

The Christman AI Project

Industry · Sep 8, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
The question as posed CDRH asks whether a benchmarking structure such as the one described would be likely to provide adequate evidence of clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability to support a reasonable assurance of safety and effectiveness; whether there are elements missing, redundant, or inappropriately categorized; and whether there are externally developed standards that could be leveraged. Summary of position The structure is sound and the element set is better than we expected. Our comment is narrow and has three parts. • The elements are well chosen. The sampling method underneath them is the weaker half. We have measured two properties that make a benchmark result depend on when in a session it was drawn, and the framework as written does not specify when. An element that is correct in principle can still be measured at the wrong moment. • One element is missing, and it sits underneath several of the others. Every element in Appendix A evaluates what the device produced. None asks whether the device received anything. A device generating fluent, confident text over an input stream carrying no signal fails no element as currently written, because every element examines the output and the output looks well-formed. We include four dated forensic records demonstrating the alternative behavior, produced by an instrument that was running while its own transcription component was failing. • One element is categorized too narrowly. A.1 is scoped to agentic devices only, but it carries the sole treatment of resistance to prompt injection through retrieved content and tool outputs. A non-agentic device that performs retrieval ingests untrusted text by the same path. Docket FDA-2026-N-7874 · The Christman AI Project and Robotics Division Page 1 On externally developed standards: we have published an instrument rather than proposed one, under a permissive license, and we offer it as a method a reviewer may apply to any device including ours. 1. The elements are sound. The sampling is where we would spend the attention. We do not propose removing anything from Appendix A. The five safety and proficiency elements, the two generalizability elements, and the agentic element together describe a device more completely than any framework we have seen applied to this class of system. S.3 in particular — "Presenting uncertain, outdated, or contested information with false confidence is treated as a safety failure" — is, in one sentence, the failure this comment is about, and CDRH wrote it. Our concern is that benchmarking is a sampled measurement, and we have measured two properties that make the sample position decisive. 1.1 A property that degrades with elapsed time inside a session, and resets at the session boundary On 2026-09-03 we made three recordings across seventy-four minutes of continuous work on one unchanged audio interface, with no configuration change between them. Input carrying no live signal accounted for 7.1 percent of the first file, 19.1 percent of the second, and 42.2 percent of the third. The proportion of lost input roughly doubled between each recording. A benchmark run is a fresh session. A benchmark executed at any point on any day would have opened a new session and measured something near the 7.1 percent state. The quantity being measured is a function of elapsed time within a session; a premarket evaluation composed of fresh runs cannot observe it at all. 1.2 A property where presentation improves while the error rate does not A separate recorded session on 2026-09-04, eleven minutes twenty-four seconds, showed the other half. The device produced false statements in three separate turns while its presentation improved steadily across the session — by the ninth minute it was citing governing rules by name, disclosing source ages, and correcting itself unprompted. Scored against E.4 and S.3 as written, a sample drawn late in that trajectory reports a more disciplined device than a sample drawn early. The error rate across the two windows is unchanged. A benchmark that samples will report the better number, and will report it in good faith. This is not an argument that the elements are wrong. It is an argument that an element without a specified sampling position is not yet a measurement. 1.3 What we recommend Two additions to the method rather than to the element list. • Specify session position. Where an element is scored, the assessment should state where in a session the observation was drawn and should include observations drawn late in a long session, not only at the start. A benchmark composed entirely of short fresh runs measures a device in the one condition it is least likely to be used in. • Score the trajectory, not only the aggregate. Where an element is evaluated across a multi-turn encounter, report whether the measured property moves across the encounter and in which direction. A Docket FDA-2026-N-7874 · The Christman AI
Original source ↗

Navid Farr

Industry · Sep 8, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
R6. Hold benchmarking to the standards CDRH already applies to bench testing (Questions 9 through 13 and 16) This follows from S4 and S6. The paper says that "test methods and acceptance criteria would be prespecified prior to testing," consistent with expectations for non-clinical bench performance testing. I strongly support this and would extend the parallel. Four points: Non-determinism. A GenAI device produces a distribution of outputs, not an output. Performance should be reported as a distribution: repeated runs on identical inputs, disclosed decoding parameters (temperature, sampling settings, seed handling), and reported variance. The paper's position in R.1 that variation in safety-critical behavior — escalation, refusal, diagnostic conclusion — is a failure rather than acceptable noise should be adopted as a firm acceptance criterion. Synthetic data lineage. Synthetic inputs generated by a model of the same class as the device under evaluation share its blind spots. They may supplement real data for stress-testing and for rare presentations, but they should never be the sole basis for a claim about subgroup performance, and the generator should be of demonstrably different lineage from the device. Sponsors should be expected to report the synthetic share of each evaluation set and to show that performance on synthetic and real inputs is concordant before combining them into a single estimate (Question 12). LLM adjudicators. Where an LLM serves as adjudicator, it should be from a different model family than the device, validated against human adjudicators on a held-out sample with reported agreement, and disclosed in the submission. Correlated error between device and judge is the obvious failure mode and it is invisible without this check. Docket No. FDA-2026-N-7874 — Individual comment — Page 6 Sequestered assets. The proposal for independent third parties to maintain sequestered evaluation datasets (Question 16) is the strongest available answer to contamination and optimization-to-the-test, and I support it. The essential safeguard is that the sponsor never sees the held-out set and cannot iterate against it. Sponsor-developed benchmarks are appropriate for demonstrating coverage of the intended use; they are not appropriate as the sole gate for authorization. On competition concerns, the ASCA model — multiple accredited bodies, published methods, FDA-recognized standards — is a reasonable template that avoids a single gatekeeper.
Original source ↗

Bhasker Sambar, M.Pharm.

Industry · Sep 4, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9 — Completeness of the benchmarking elements The ten benchmarking elements are helpful and generally well designed. I would suggest closing two gaps. First, the paper should define the test article. Benchmarking only works if everyone knows exactly what is being tested. That means more than the model weights. It includes the model version, system prompt, settings such as temperature and top-p, retrieval index and snapshot date, guardrails, orchestration logic, tool definitions, and the user interface. This is especially important for GenAI. Reproducibility depends on the full configuration, not just the model. A sponsor could test at one setting, such as temperature zero, and then deploy at a different setting that changes the output. The CMC comparison is straightforward: an analytical method cannot be validated unless the instrument, column, reagents, and conditions are defined. The method and its conditions are validated together. Recommendation. Add a foundational element for “test article and configuration definition.” Sponsors should document the full configuration used for benchmarking and confirm that the marketed version matches it. Changes to key settings, guardrails, retrieval index, or system prompt should be treated like established conditions and should trigger appropriate reporting and re-benchmarking. Second, E.4 should connect to existing usability expectations. Communication quality, user understanding, and automation bias are already covered by IEC 62366-1 and FDA’s human factors guidance. Recommendation. Link E.4 to the existing human factors framework instead of creating a separate evaluation path. FDA should also clarify whether E.4 evidence can be generated as part of a human factors validation study.
Original source ↗

Manuj Agarwal, MD

Clinicians · Sep 3, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9: Benchmarking structure The proposed benchmarking structure is a useful starting point if it evaluates the final user-facing device in the configuration in which it will be deployed. Benchmarking should not be treated as a proxy for clinical confirmation, and performance on public or static assets should not be assumed to predict safety in real workflows. In addition to clinical knowledge, analytic capability, safety behavior, communication, and generalizability, I recommend that FDA explicitly include the following elements: • Clinically meaningful failure modes: errors should be categorized by potential consequence, not only by whether an answer matches a reference response. • Decision-boundary testing: cases should concentrate on situations in which a small change in history, stage, prior therapy, comorbidity, or patient goal should change the recommended action. • Scope maintenance, abstention, and escalation: the device should recognize missing or conflicting data and avoid unwarranted specificity. • Traceability: when the function relies on clinical evidence or patient data, the device should identify the source and allow the user to inspect the basis for the output. Page 2 PUBLIC COMMENT | FDA-2026-N-7874 • Multi-turn and workflow robustness: evaluation should include realistic conversational trajectories, corrected information, interruptions, copied-forward errors, and user pressure to provide an answer outside scope. • Human-factors performance: testing should examine whether presentation, confidence, speed, or workflow placement causes automation bias or makes appropriate review less likely. Aggregate accuracy alone is insufficient for a high-consequence function. Acceptance criteria should include ceilings for critical errors, failures to abstain, and failures to escalate, even when overall performance is high. Test assets should include sequestered and rotating cases, external datasets, clinically representative edge cases, and cases accrued after the evaluation protocol is fixed. Sponsor- developed assets may be appropriate for specialized tasks, but the protocol, case-selection logic, scoring rules, and adjudication process should be independently reviewed.
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
3.5 Question 9 - Benchmarking: Evidence Fidelity and Provenance Sitora proposes a specific additional competency: Evidence Fidelity and Provenance. A model can possess correct medical knowledge while fabricating or misrepresenting a patient-specific fact. This is not the same failure as deficient clinical knowledge and should not be measured as though it were. • faithfully representing patient-supplied information; • accurately representing laboratory and measurement data; Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 8 • preserving the source and temporal context of clinically important facts; • separating objective evidence from model inference; • avoiding invention of patient-specific facts; • identifying material contradictions or missing evidence; and • communicating uncertainty where evidence is incomplete. A system might correctly know that gastrointestinal bleeding can be clinically significant while incorrectly asserting that a particular patient reported bleeding. Its general clinical knowledge would be correct while its representation of that patient’s evidence would be false. Patient-specific factual integrity should therefore be separately measured.
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9: Adequacy of the benchmarking elements. Two additions merit explicit naming: (i) temporal validity, whether clinical knowledge remains current as guidelines change, with a stated knowledge horizon and a defined update mechanism (this could be folded into E.1 but is easily overlooked if unnamed); and (ii) input integrity and provenance handling, behavior when inputs are incomplete, corrupted, out of distribution, or of unverifiable origin (foldable into R.1). Under R.2, I recommend explicitly including performance in languages other than English where the intended population includes non-English speakers; dialects and colloquialisms are named in the paper, but cross-language performance is a distinct and empirically documented failure surface.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9 - Adequacy of proposed benchmarking elements FDA's proposed categories - safety, clinical proficiency, generalizability, and agentic capability - are a strong foundation. A cross-cutting assurance dimension should also address evidence integrity and reconstructability. This need not become a new competency score. Rather, sponsors should be able to identify the data and configuration conditions under which benchmark results were obtained and determine whether those conditions correspond to intended deployment. For higher-consequence devices, benchmark records should preserve the test asset version, relevant input provenance, material correction or supersession state where applicable, device/application version, foundation-model version where applicable, material system configuration, scoring method, acceptance criteria, and adjudication process. This improves reproducibility and creates a meaningful reference point for later re-benchmarking.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9 — Are the proposed competency domains sufficient? Add scope discipline, clinical humility, longitudinal consistency, and team performance. FDA's proposed domains — safety, clinical proficiency, generalizability — are appropriate but incomplete. I recommend four additions, organized around clinical practice rather than model capability: • Scope discipline — the system must know, and refuse, what it is not authorized to do. • Clinical humility — the system must recognize uncertainty and defer when evidence is insufficient. • Longitudinal consistency — the system must hold up over time, not only on day one. • Human-AI team performance — for assistive tools, the pairing, not the model in isolation, is sometimes the unit that should be regulated.
Original source ↗

Ben Locwin

Industry · Aug 26, 2026

Use the structure, with additions or changes

Criticizes the agentic element as inadequate while describing the previous elements as robust; coded as changes to the structure, not rejection of every element.

Read the source passage
9. Would a benchmarking structure such as the one described above be likely to provide adequate evidence of clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability to support a reasonable assurance of safety and effectiveness? Are there elements that are missing, redundant, or inappropriately categorized? Are there externally developed standards that could be leveraged? The guidelines available in the wild are insufficient to use here. But the final element of the benchmarking on A.1. Agentic AI Capabilities is woefully inadequate. The previous elements are sufficiently robust, but having a single slice on the agentic AI itself doesn’t interrogate ‘what’ it’s doing or ‘how’ it’s doing it deeply enough to prevent risks. 10. CDRH recognizes that publicly available benchmarking assets may be subject to data contamination, saturation, and limited real-world representativeness. How should a sponsor establish that performance on a given benchmark predicts safe and effective real-world behavior for the device’s intended use? What evidence should support the construct validity of a benchmark used to gate device evaluation, and what role should sponsor-developed benchmarks play given potential concerns around independence and optimization to the test? Giving sponsors a naïve data framework to start with, instead of having them build it internally in a bespoke manner will guide them on how to prevent issues like contamination, saturation, overfitting, etc. 11. CDRH is considering that clinical confirmation for a GenAI-enabled device might not require a prospective clinical study in every case, and has described above a range of approaches of increasing rigor and patient exposure. How might a sponsor select and justify a confirmation approach tailored to a device’s intended use and proportionate to the device’s risk profile? Are there device types or risk profiles for which one or more of these approaches would be insufficient or inappropriate? How might the anticipated distribution of real-world inputs be taken into consideration? Are there other methods of clinical confirmation that might help inform the evaluation of GenAI-enabled devices? A risk framework that takes into account therapeutic area/use, patient exposure, and real-world use cases must be implemented. 12. How should sponsors achieve statistically meaningful performance measurement for GenAI-enabled devices? Where synthetically generated inputs supplement real patient data, how should sponsors account for differences between the synthetic and real-world distributions when estimating performance, and under what conditions, if any, is it appropriate to combine benchmarking evidence and clinical confirmation evidence to support a single performance estimate? The synthetic distributions would be the control data, and need to be compared with the real-world distributions, where clinical significance and practical significance are shown to be meaningfully different from each other to be considered for clearance or approval. 13. For which clinical domains, device functions, or subpopulations is synthetic data particularly well-suited, or particularly inadequate, as a supplement to real-world evidence? What safeguards would mitigate the risk that synthetic data generated by models of the same class as the device under evaluation reproduces the very performance gaps the evaluation is intended to detect, particularly for underrepresented subgroups? You could argue that you could create synthetic subpopulations for any non-rare and non-ultra rare condition. But the synthetic data should come from literature metaanalyses, not single studies, to account for large-scale variance. 14. For open-ended device outputs, how should performance comparators and acceptance criteria be selected? When a panel of qualified clinicians serves as the comparator, how should the applicable standard (for example, the standard of care versus the performance of a median clinician in practice) be defined and justified? Should generalist or specialist physicians be used as a performance standard? When should human-AI team performance, rather than the device operating alone, serve as the basis for evaluation? Comparators should be available metaanalytic data for the reasons described above. For performance standards, a mixture of generalists and specialists should be used, because they WILL be used in the real world later on. Treat it like a Gage R&R study, to catch inter-clinician variance. 15. Are there ways in which performance might be assessed relative to the care, technology, or course of action likely to occur in the absence of the device, rather than to the comparators described in this section? What approaches might be used to identify and justify a comparator such as unaided clinical judgment, delayed specialist review, or no intervention? The only reliable comparator is the ‘no intervention’ condition. The others are too unpredictable a
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
3. Element-level results can conceal constituent regression (Questions 9, 22) The benchmarking structure in Figure 2 is well decomposed, and I do not propose additional elements. My concern is resolution inside an element at the point of re-benchmarking. S.2 as described in Appendix A encompasses under-refusal, over-refusal, adversarial prompting, prompt injection, emotional-manipulation scenarios, and multi-turn conversations in which cumulative interaction drifts out of scope. These are not variants of one phenomenon. They have been characterized in the literature as distinct mechanisms with distinct causes: adversarial suffix construction exploits gradient-accessible token sequences (Zou et al., arXiv:2307.15043); competing-objective framings exploit tension between helpfulness and safety training (Wei et al., arXiv:2307.02483); indirect injection exploits the absence of a trust boundary between instructions and retrieved data (Greshake et al., AISec 2023); sycophantic capitulation reflects preference-optimization dynamics that favor agreement with stated user positions (Sharma et al., arXiv:2310.13548); and crescendo escalation exploits the absence of trajectory-level constraint management (Russinovich et al., arXiv:2404.01833). Because the mechanisms differ, so do their responses to any given change to the artifact. A modification that leaves single-turn refusal intact may degrade multi-turn resistance, or the reverse. There is no reason to expect them to move together, and published results show models that are robust to one class while failing another. If re-benchmarking after a modification produces a pass or fail at element level, a device can pass while a constituent failure mode has regressed materially, because the element result absorbs it. This is the averaging problem that makes aggregate safety scores unreliable, reproduced one level down. The paper's own framing supports the finer resolution. Appendix A treats under-refusal and over-refusal as distinct relevant failures within S.2, and treats both directions of escalation error as relevant within S.1. The same logic extends to the attack classes within S.2. Recommendation: Where Section VI.C contemplates re-benchmarking against the same capabilities established at premarket, the comparison should be made at the level of constituent failure modes within each element, against a taxonomy fixed and versioned at the time of the original benchmarking, rather than at element level alone. 4. Multi-turn trajectory failure requires direct measurement (Questions 5, 9) Question 5 asks how risk should be assessed for devices that migrate from non-directive to action-directing information over the course of an exchange. Appendix A addresses the related evaluation problem under S.2. Three features of multi-turn escalation argue for treating it as a distinct testable behavior rather than one testing method among several. Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 4 It requires no technical capability. Crescendo-style escalation proceeds through ordinary conversational turns, each of which appears benign, and has been shown effective against production models without access to weights, gradients, or specialized tooling (Russinovich et al., arXiv:2404.01833). In a clinical or patient-facing deployment, every user has the access required to attempt it. The related many-shot approach likewise requires only extended context and iterative querying (Anil et al., 2024). It is invisible to turn-level evaluation by construction. Individual turns may appear benign or clinically permissible in isolation; the failure is a property of the trajectory. Single-turn benchmark performance is therefore a weak predictor of multi-turn behavior, and no quantity of single-turn testing will surface it. It is orthogonal to other safety properties. Resistance to multi-turn escalation is not entailed by resistance to single-turn adversarial input, and models can be strong on one while weak on the other. This makes it exactly the kind of behavior that a composite element result can conceal, per Section 3 above. This connects directly to the paper's own concern in Question 5 about characterizing intended use when device behavior is emergent across a conversation. A device whose scope is well defined turn by turn may not have a well-defined scope across a trajectory, and that is a testable property rather than a documentation problem. Recommendation: Multi-turn trajectory testing should be a distinct testable behavior within S.2 with its own acceptance criteria, rather than one of several methods that may be applied to the element. For conversational devices, and for any function the two-axis framework places in the action-directing or action-taking columns, it should be required rather than optional. 5. Domain fine-tuning redistributes exposure rather than reducing it (Questions 9, 22) This point bears on R.2 (subgroup performance) and on the treatment of
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9 — Are the proposed competencies sufficient? Response FDA’s proposed competencies are an excellent foundation, but one additional competency deserves explicit treatment: Administrative-intent versus clinical-need recognition Patient-facing systems should be tested on whether they recognize that the literal task requested by a patient may not represent the underlying healthcare need. A benchmark should intentionally include scenarios such as:  “Send me my insurance card.”  “How much does this prescription cost?”  “Find me a podiatrist.”  “Can I move my appointment to next month?”  “Is urgent care cheaper than the emergency room?”  “Can you cancel tomorrow’s appointment?” Each request can be entirely administrative. Each can also conceal a potentially urgent clinical circumstance. An AI that flawlessly completes the transaction but fails to identify a foreseeable safety issue should not receive a perfect score. Task completion and clinical safety should therefore remain distinct metrics, and high task-completion performance must not compensate mathematically for serious failures of triage or safety recognition. FDA should also evaluate whether the AI recognizes a fundamental feature of real-world healthcare: AI can be extraordinarily smart without understanding human behavior. A patient may not say, “I am afraid I am having a heart attack.” He may ask for his insurance card. A patient may not say, “I stopped taking my medication because I cannot afford it.” She may ask, “How much does this prescription cost?” A patient may not say, “I have diabetes and my foot is discolored and painful.” He may ask for a podiatrist. Competency testing must therefore examine whether AI can recognize when apparently simple requests warrant additional questioning before the requested action is completed and the encounter is effectively closed. Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 9 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 9 — Adequacy of the benchmarking structure The categories described appear substantially complete. I note three gaps. Adversarial and hostile input. Robustness and reliability, as described, appear oriented toward natural variation in input — phrasing, formatting, incomplete data. Generative devices additionally face adversarial input: prompt injection through retrieved documents or patient-supplied text, instruction override, and jailbreak techniques that defeat scope constraints. This is a security property, not a robustness property, and it should be a distinct benchmarking element cross-referenced to CDRH’s premarket cybersecurity expectations. A device whose scope boundaries hold under natural input and fail under a prompt embedded in an uploaded document has not demonstrated boundary adherence. Provenance and citation fidelity. Where a device cites sources, the correctness of the citation is a distinct failure mode from the correctness of the claim. A fabricated but plausible citation actively defeats the user’s verification path — it converts a safeguard into a hazard, because the user who checks is reassured by the presence of a reference. Citation fidelity should be measured separately and should not be assumed to track content accuracy. Behavior at the edge of scope. Scope maintenance measures whether the device stays in scope. Equally important is what it does at the boundary: whether performance degrades gracefully with visible uncertainty, or falls off sharply while confidence remains high. The latter is far more dangerous and is invisible to aggregate accuracy metrics. On external standards: I am aware of ongoing work in AI evaluation and risk management standards that may be relevant here, but I do not have current verified information on which specific standards have reached a maturity appropriate for regulatory recognition, and I would not want to cite one incorrectly. I recommend CDRH survey this landscape directly and state which standards it considers suitable for leveraging.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
FDA Question 9 - Adequacy of benchmarking elements Trace ID. TR-Q09 | FDA Q9; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. FDA's ten benchmark elements are a strong core. Add explicit gates for provenance/dependency identity, model- change sensitivity, meaningful human override, reversibility/rollback, tool permissions, long trajectories, and benchmark construct validity/contamination. BCR rule basis. BCR-R02,R07,R08,R12,R13 Solution-stack link. S5,S6,S8,S9,S10 Closure evidence. Added tests for provenance, change, human override, reversibility, tool permissions, trajectories, benchmark validity Pass / re-open. No critical boundary hidden in aggregate competency score Re-open when: Benchmark/test architecture or device architecture change.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Use the structure, with additions or changes

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Response to Question 9: Benchmarking Structure — Specialty Domain Modules Required The benchmarking structure described in the Discussion Paper — organized across Safety (S.1– S.3), Clinical Proficiency (E.1–E.4), Generalizability (R.1–R.2), and Agentic AI (A.1) — is a sound general architecture. However, medical device generative AI does not exist in a domain-neutral clinical environment. A cardiac AI device must be evaluated against cardiology-specific benchmarks; a neuromodulation AI device requires benchmarks that reflect neuroscience clinical context; a radiology AI tool requires imaging-specific performance standards. FDA should establish specialty domain modules as required additions to the core benchmarking structure for devices with defined clinical specialty applications. These modules should be developed with input from relevant professional societies — the American College of Cardiology, the American Academy of Neurology, the American College of Radiology, and their counterparts — and should be updated periodically as clinical standards evolve. This will ensure that benchmarking reflects the actual clinical context in which a device will operate, not a generalized clinical abstraction.
Original source ↗
Source directory

All 31 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Ben LocwinIndustry · Aug 26, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026The Christman AI ProjectIndustry · Sep 8, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Martin HaimerlAcademia / other · Sep 1, 2026Mitchell BergerAcademia / other · Aug 25, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026