← All 104 filings

Ben Hyams

Awaiting reviewAwaiting reviewFiled September 20, 20261,262 words · 1 attachmentFDA-2026-N-7874-0107
Not yet read. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

See attached file(s)

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

PUBLIC COMMENT TO THE U.S. FOOD
AND DRUG ADMINISTRATION
Considerations for the Regulation of Generative AI-Enabled
Medical Devices
Discussion Paper and Request for Feedback

Docket No. FDA-2026-N-7874

Submitted by: Ben Hyams, BA
Date: Sept 20, 2026

This comment responds to the U.S. Food and Drug Administration’s August 2026 Center for
Devices and Radiological Health (CDRH) discussion paper, Considerations for the Regulation of
Generative AI-Enabled Medical Devices, and request for feedback.1 This response is offered
from a clinical perspective, with particular focus on standards for premarket assessment and on
ensuring safety, effectiveness, and value for patients. The FDA’s discussion paper specifically
seeks input on risk assessment, premarket evaluation, and postmarket monitoring of GenAIenabled medical devices.

Generative AI (GenAI) devices are likely to substantially expand the role of medical devices.
The broad functionality of GenAI systems—including autonomous decision-making, actiondirecting, and action-taking capabilities—creates new challenges that the current regulatory
paradigm was not designed to address. As CDRH develops new approaches for these
technologies, it is essential that the goals of applying “least burdensome” principles and
facilitating timely access do not displace FDA’s core mandate to provide reasonable assurance of
device safety and effectiveness.

It is worth noting that current standards of evidence for device clearance also present challenges
for traditional, non-GenAI devices. The 510(k) pathway, the most commonly used medicaldevice marketing pathway, does not generally require clinical testing when substantial
equivalence can be demonstrated. Over successive generations of devices, lenient interpretations
of substantial equivalence can permit substantial changes in technology or scope without
corresponding clinical evaluation. At the same time, nonclinical validation can rely heavily on
surrogate or intermediate measures and on testing datasets that may have limited clinical
generalizability. The development of a regulatory framework for GenAI therefore presents an
opportunity not only to address the unique risks of GenAI, but also to strengthen standards for
evidence for the next generation of medical devices.

CDRH’s discussion paper proposes a premarket evaluation approach based on “competency
assessment,” including nonclinical device benchmarking and clinical confirmation to assess
whether a GenAI-enabled device performs as intended. The paper appropriately recognizes the
potential need for clinical confirmation while also positing that clinical confirmation might not
require a prospective clinical study. I believe prospective clinical studies should be the principal
mechanism for establishing clinical safety and effectiveness when a device directly influences
patient care. Benchmark testing and simulated clinical scenarios can provide valuable evidence,
but they have several important limitations:

1. Human interaction with medical devices is inherently variable.
Idiosyncratic characteristics of patients and users are unlikely to be fully represented in
benchmark datasets. The clinical behavior of a GenAI system may change depending on the
intricacies of human language and psychology such as how patients describe symptoms, what
information they volunteer, and how users respond to recommendations.

2. Users may interpret GenAI outputs differently from expert evaluators.
An output judged to be technically correct by an expert assessor may be misunderstood by a reallife patient or clinician. For example, language that is technically accurate but heavy in medical
jargon may lead to an incorrect clinical action.

3. Device performance depends on more than the generated output.
Other components of the system including the user interface, workflow, and system-level
interactions may materially affect how an output is understood and acted upon. The FDA’s
existing multiple function framework recognizes the need to assess safety and effectiveness of
multiple functions on an integrated, systems level, which may not be captured by benchmark
testing.2

4. Some clinically important risks may be difficult to anticipate prospectively.
GenAI systems can exhibit emergent behaviors that may not become apparent in benchmark
testing. For example, recent research has demonstrated that optimizing language models for
helpfulness can increase the tendency to reinforce misconceptions among patients, and result in
potentially harmful information.3 Such findings illustrate the possibility that clinically relevant
risks may emerge from the interaction between model behavior and human users.

The benchmarking methodology contemplated by CDRH raises a broader methodological
concern: qualitative assessment of GenAI outputs by human or machine-based evaluators may
function as an intermediate or surrogate measure of clinical performance. Appropriate model
outputs are presumed to correlate with downstream clinical outcomes, but that relationship may
not always be established. Moreover, the necessary reliance on a limited test set and the
possibility of data contamination both pose threats to the generalizability of results. Benchmark
performance should therefore be interpreted cautiously, particularly for devices that directly
influence diagnosis, treatment, or other consequential clinical decisions.

The broad functionality of some GenAI devices may make it more difficult to design prospective
clinical studies that evaluate every possible parameter of safety and effectiveness. However, this
challenge should not be treated as a reason to abandon prospective clinical evaluation where
meaningful clinical endpoints are available.
Some GenAI devices target discrete clinical outcomes that can be assessed through relatively
straightforward clinical studies. Consider, for example, a patient-facing conversational device
that provides insulin-dosing instructions based on patient-reported meals, symptoms, or
adherence. The clinical effectiveness of such a device is better captured by outcomes such as
glycemic control or medication adherence than by whether an expert evaluator considers the
device’s textual responses appropriate. For such a device, an appropriate premarket standard
could include a well-designed clinical study demonstrating non-inferiority on prespecified
clinically meaningful outcomes, together with appropriate safety endpoints.4

Other GenAI devices, such as a virtual primary care provider, may have a scope too broad to be
captured by a small number of discrete clinical variables. In these cases, a complementary
evaluation strategy may be appropriate. One component could consist of clinically meaningful
patient outcomes (e.g. disease-specific biomarkers, emergency-department visits, validated
patient-reported outcomes), while a second component could involve independent clinician
adjudication of real patient cases. Retrospective analysis or prospective “shadow” deployment
may also provide useful evidence in limited applications where the GenAI system and clinicians
receive identical inputs. However, these approaches are unlikely to fully reproduce the
interactive nature of patient-facing GenAI systems.

At this stage of GenAI device development, when the full spectrum of clinically relevant risks
remains uncertain, the threshold for prospective clinical evaluation should be relatively low for
devices that directly influence patient management. Devices that engage in action-directing or
action-taking activities should generally receive prospective clinical scrutiny, regardless of the
anticipated consequence level.

Benchmarking and clinical simulation remain valuable tools. They can support research and
development, characterize model behavior, monitor changes in performance over time.
Benchmarking and clinical simulation could potentially support clearance of lower-risk, nondirectional, informational devices. They should not, however, be assumed to establish clinical
safety and effectiveness when a device's intended use directly affects patient care.

Ultimately, evaluation of GenAI-enabled medical devices should remain grounded in patientcentered clinical outcomes. Greater reliance on benchmark performance, surrogate measures,
machine-based assessment creates a risk that regulatory evaluation becomes increasingly
detached from the outcomes that matter most to patients.
Disclosures

The author has no conflicts of interest or relevant financial disclosures to report.

Bibliography

1. U.S. Food and Drug Administration, Center for Devices and Radiological Health.
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion
Paper and Request for Feedback. August 2026.

2. U.S. Food and Drug Administration. Multiple Function Device Products: Policy and
Considerations - Guidance for Industry and Food and Drug Administration StaP. July
2020.

3. Chen S, Gao M, Sasse K, et al. When helpfulness backfires: LLMs and the risk of false
medical information due to sycophantic behavior. Npj Digit Med. 2025;8(1):605.
doi:10.1038/s41746-025-02008-z

4. Hyams B, Kerlikowske K, Redberg RF. New Mammography Tools — The Need for
Clinically Meaningful Assessment Standards. N Engl J Med. 2025;393(3):211-213.
doi:10.1056/NEJMp2500274