← All 104 filings

Denise Bowen

Awaiting reviewAwaiting reviewFiled September 20, 20261,064 words · 1 attachmentFDA-2026-N-7874-0106

FDA questions it names

Questions this filing names by number.

Q7 · The competency-based approachQ10 · Benchmark contamination and saturationQ20 · Machine-based supervisory agentsQ25 · Foundation Model Master Files

Not yet read. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Please see the attached document for my full comment. I address Questions 7, 10, 20, and 25 of the discussion paper, drawing on peer-reviewed clinical AI literature to evaluate the two-axis risk framework, the competency-based evaluation approach, postmarket supervisory-agent monitoring, and the proposed Foundation Model Master File program.

Thank you for considering these comments.
Denise Bowen
Founder & CEO, DB Connect | Advisor, UN Office on Drugs and Crime
Submitted in an individual capacity

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

PUBLIC COMMENT TO THE U.S. FOOD AND DRUG ADMINISTRATION
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper
and Request for Feedback

Docket: FDA-2026-N-7874
Submitted by: Denise Bowen

Professional background: Founder and CEO, DB Connect; Advisor, UN Office on Drugs and Crime

Capacity: Individual comment; views expressed are my own

Date: September 19, 2026

Questions addressed: 7, 10, 20, 25

To the Food and Drug Administration:

Thank you for the opportunity to comment on the Center for Devices and Radiological Health's
discussion paper on the regulation of generative AI-enabled medical devices.

I am the Founder and CEO of DB Connect. For the United Nations Office on Drugs and Crime, I advise
Chief Legal Officers and Attorneys General from UN Member States on cybercrime, artificial intelligence,
and digital governance. My work covers electronic evidence, cross-border cooperation, and implementation
of the United Nations Convention against Cybercrime. I assess the legal, regulatory, and security
implications of emerging technologies to strengthen institutional capacity and develop governance
frameworks that safeguard the rule of law while enabling responsible innovation. I also serve on the
EmblemHealth Advisory Council, shaping AI-governance standards across a payer-provider healthcare
ecosystem, and I lead DB Connect's AI Governance Committee, which applies the NIST AI Risk
Management Framework to pre-deployment bias audits. I submit this comment in my individual capacity.

Response supporting Section II (Background, p. 2)
The paper identifies confabulation, an effectively unbounded input-output space, and post-deployment
behavioral drift as defining risks of generative AI. This characterization is well supported by the current
empirical literature. Adversarial testing across six large language models found that a single fabricated
clinical detail, embedded in a prompt, was repeated or elaborated on in 50% to 82% of outputs across 5,400
generated responses (Omar et al., Communications Medicine, 2025). A separate study of 12,197 diagnostic
large-language-model outputs found accuracy ranging from 2.92% to 46.34% depending on the model
tested, with omission of applicable clinical guidelines reaching 97.08% in the weaker-performing model
(van Kessel et al., BMJ Health & Care Informatics, 2026). These findings suggest that the paper's framing
of GenAI risk is not overstated. If anything, its tone reads more confidently than the current evidence base
supports.
Response to Question 7 (narrative: Section V.A, p. 10; formally numbered: Section V.E, p.
17, and Appendix B, p. 28): Usefulness of the competency-based approach
The analogy to clinician licensure is a reasonable organizing principle that reflects live scholarly debate
rather than a novel proposal. Patel and Blumenthal argue in JAMA Health Forum (2026) that current device
frameworks are ill-equipped for generative AI. They propose a licensure-like model built on the training,
evaluation, and lifelong oversight of clinicians, a framing that matches the paper's own citation. I
recommend one addition. Clinician-panel disagreement is a documented confound in diagnostic referencestandard studies, yet Section V.D.1's discussion of comparator selection does not address it (p. 16). The
paper should require that any comparator clinician panel report its own inter-rater reliability statistics, such
as Cohen's kappa.

Response to Question 10 (narrative: Section V.B, p. 12; formally numbered: Section V.E, p.
17–18, and Appendix B, p. 28–29): Benchmark contamination and saturation
On benchmarking (Question 10), I cite peer-reviewed evidence to support a requirement that sponsors
disclose concrete safeguards against benchmark contamination and saturation. A study of six leading large
language models found that models reproduced or elaborated on adversarially planted false clinical details
in approximately 50 to 83 percent of tested outputs (Omar et al., Communications Medicine, 2025). A
separate study of 12,197 diagnostic large-language-model outputs found accuracy ranging from 2.92 to
46.34 percent depending on the model tested, with omission of applicable clinical guidelines reaching 97.08
percent in the weaker-performing model (van Kessel et al., BMJ Health & Care Informatics, 2026). The
broader benchmarking literature, including hallucination-specific benchmarks such as MedHallu and MEDHALT, documents this as a persistent, unresolved problem rather than a theoretical one. I recommend that
sponsors be required to disclose the specific mitigations used against these known failure modes.

Response to Question 20 (narrative: Section VI.A, p. 19–20; formally numbered: Section
VI.D, p. 21, and Appendix B, p. 29): Machine-based supervisory agents
The paper's discussion of supervisory agents does not address a circularity risk. An oversight model built
on similar underlying architecture may share the same failure modes as the device it supervises. In the
adversarial-hallucination study cited above, mitigation prompting reduced the mean hallucination rate from
66% to 44%, while temperature adjustment produced no measurable improvement. A supervisory model
watching a clinical model therefore retains a meaningful chance of failing independently rather than
catching the primary model's error. I recommend that supervisory agents be required to undergo the same
S.1 through S.3 safety benchmarking elements described in Section V.B before being credited as a valid
postmarket-monitoring control.

Response to Question 25 (Section VII.A): Foundation Model Master Files
Model developers have limited commercial incentive to disclose safety-relevant limitations of their own
models. I recommend that the paper's proposed Foundation Model MAF program be structured as a
mandatory disclosure requirement for higher-consequence devices, rather than a purely voluntary one. This
recommendation carries particular weight given that no generative-AI-enabled device has yet been
authorized for marketing (Congressional Research Service, IF13245, June 2026), which means every
mechanism proposed in this paper remains untested against a real postmarket record.
Conclusion
The paper accurately reflects an active and unresolved area of clinical AI evidence. Its principal gap is
context. It does not acknowledge that the entire regulatory framework it proposes has no operating
precedent, since no generative-AI-enabled device has yet cleared FDA authorization. I encourage FDA to
treat the two-axis risk framework, together with the competency-based approach, as a sound starting
scaffold. FDA should also require statistical reliability reporting for clinician comparators, safety
benchmarking for any AI used to supervise other AI, and mandatory rather than voluntary foundationmodel disclosure.

Thank you for considering these comments.

Denise Bowen
Founder & CEO, DB Connect | Advisor, UN Office on Drugs and Crime
Submitted in an individual capacity