UltraAI (Reza Rahmanzadeh)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
UltraAI respectfully submits the attached comments on the discussion paper "Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback" (Docket No. FDA-2026-N-7874). Please see the enclosed document for our full comments. A one-page figure of our proposed framework appears on page 1, and a summary of our comments and recommendations on pages 3–6.
Reza Rahmanzadeh, MD, PhD
CEO & Founder @ UltraAI
Attachment
Comments of UltraAI on the FDA Discussion Paper
“Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion
Paper and Request for Feedback”
DOCKET NO. SUBMITTED BY AUTHOR DATE
FDA-2026-N-7874 UltraAI Reza Rahmanzadeh, MD, PhD 28 September 2026
Docket No. FDA-2026-N-7874 · Contents
Contents
I Introduction and Statement of Interest........................................................................................... 2
II Summary of Comments and Recommendations........................................................................... 3
III Comments on Section IV — Considerations for the Assessment of Risk ............................................ 7
Questions 1 through 6
IV Comments on Section V — A Competency-Based Approach for Premarket Evaluation ....14
Questions 7 through 17
V Comments on Section VI — Postmarket Monitoring ............................................................................................. 21
Questions 18 through 20, 24 and 25
VI Architectural Safeguards for High-Risk-Profile Device Functions, Including Autonomous
Functions...................................................................................................................................................................................................... 25
Questions 4, 11, 18 and 26
VII Applicability Beyond Generative Architectures ........................................................................................................ 28
Questions 1, 17 and 26
VIII Generalizability: Synthetic Data and Independent Benchmarking .......................................................... 30
Questions 10, 12, 13 and 16
IX Conclusion....................................................................................................................................... 34
APPENDICES
A Generating and Presenting: Illustrative Examples....................................................................... 36
B Risk Profiles: Illustrative Examples................................................................................................ 39
C Bounding and Claim-Level Evaluation: Illustrative Examples..................................................... 41
References...................................................................................................................................... 43
I. Introduction and Statement of Interest
UltraAI develops AI-enabled software for medical imaging, including radiology and pathology. The
company holds an FDA Breakthrough Device Designation and participates in CDRH’s Total Product Life
Cycle Advisory Program.
I write as a neuroradiologist and AI engineer. My doctoral research was in quantitative neuroimaging
and my postdoctoral research in AI; I design and validate machine learning systems for medical
imaging. Having worked hands-on with both conventional and generative AI tools in clinical imaging, I
comment as a clinician who relies on these tools and as an AI engineer accountable for how they
behave when they are wrong.
CDRH has asked the right questions, and the framework that emerges will shape not only what is
permitted but what is built. We share the Agency’s objective of timely patient access to safe and
effective devices, consistent with least burdensome principles. Our recommendations aim to make the
framework more precise: several propose tests that are narrower, more auditable, and less costly to
apply than the constructs currently described, while providing clearer assurance of safety.
Quotations in this comment are drawn from the discussion paper unless otherwise indicated. 1
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 2 of 45
Docket No. FDA-2026-N-7874 · Contents
II. Summary of Comments and Recommendations
CDRH has identified two regulatory challenges for GenAI-enabled devices: a risk-based approach to
classifying them and determining their requirements, and the types of valid scientific evidence needed
to evaluate them. We make three central proposals. For risk, ask whether a function generates a clinical
claim or presents one that originated elsewhere, and who receives it. For evidence, score the clinical
claims within open-ended output against a reference standard; no new category of evidence is needed.
And treat open-endedness and output variability — what makes GenAI hard to evaluate — as largely
design choices: bound them where possible, and require architectural safeguards where risk is high.
A. The 2rst challenge: a risk-based approach (Section III; Questions 1–6)
Classify functions by what their output asserts, not how it is worded. The paper’s qualifications point
to this: directiveness may depend on substance, not solely on words such as “recommend”; a “talk to
your doctor” disclaimer may not make an output less directive; and measurement functions, though
non-directive, may carry higher risk. Even the paper’s own example of non-directive output, a
cardiovascular risk score, prompts action when very high. How directive an output sounds is a poor
guide to how much it matters. We propose a test that, like the IMDRF framework FDA adopted in 2017,
grades what an output contributes to the clinical decision, and that formalizes the traceability to primary
sources Question 1 invites. A function presents a clinical claim — one concerning detection, diagnosis,
or clinical management — when the claim originates elsewhere: in its input, such as a physician’s
report, or in an identified source, such as a published guideline. It generates one when the claim
originates with the model.
The test turns on where a claim originates — not on how it is phrased, nor even on whether it is correct.
Raw data is not a claim, so a diagnosis produced from an image is always generated. A function also
generates when it goes beyond a sourced claim: changing its certainty (restating “cannot exclude
ischemia” as “ischemia is present”), applying it to a patient its source excludes, or combining sourced
findings into a diagnosis that no published rule specifies. Nor does a citation make a claim presented: a
model can attach a genuine reference to a claim it produced itself. Presenting functions should then be
benchmarked for faithfulness to their source, so that the classification is demonstrated, not declared.
Assess risk along four dimensions, which together set the evidence and safeguards required:
• Significance of the output: presenting, or generating a measurement, a detection, a diagnosis, or
a management determination, in increasing order.
• State of the condition: the paper’s consequence axis, graded critical, serious, or non-serious as
in IMDRF.
• The user: whether the recipient can judge the output, evaluate how it was produced, and act on it
within the established clinical pathway. It changes the likelihood that an error is caught, not its
severity, so it belongs in its own dimension rather than on the consequence axis. Autonomy is
part of this dimension, not a separate one: a function operates autonomously when its user is no
longer the clinician who would ordinarily perform its task.
• Characterizability: whether output is deterministic or stochastic, and inputs and outputs bounded
or open-ended — design properties that determine how much evidence a function needs.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 3 of 45
Docket No. FDA-2026-N-7874 · Contents
B. The second challenge: valid scienti2c evidence (Section IV; Questions 7–17)
GenAI does not require a new category of evidence: evaluate the clinical claims in its output. The
paper observes that for open-ended output “a single correct response often does not exist.” That is
true of the text, but not of the clinical claims the text makes. A statement that a finding is present is right
or wrong against a reference standard, and its sensitivity and specificity can be measured in wellcontrolled investigations, as 21 CFR 860.7 contemplates. We recommend that higher-risk generating
functions emit their clinical claims in structured form alongside any free text, so they can be scored
directly, without a second model judging what was asserted; where claims are instead extracted from
free text, the extractor must itself be validated. Variability is then measured where it matters: whether
the same claim recurs across repeated runs and clinically irrelevant input changes, not whether its
wording does.
Treat open-endedness and variability as design choices, and make bounding the primary
expectation. The competency-based approach and the heavier reliance on postmarket monitoring both
respond to open-endedness and output variability — properties largely within a sponsor’s control.
Inputs can be bounded by pre-specified prompts, admissibility gates, and curated retrieval; outputs by
pre-specified claim sets, report templates, and constrained decoding. Since these methods guarantee
structure, not correctness, bounding must itself be validated. Bounded devices should bear a lighter
burden; a sponsor proposing competency-based evaluation should justify why bounding is not
achievable.
Recognize that competency inference is least reliable where consequences are greatest. Licensingstyle benchmarks approach saturation while practice-oriented diagnostic tasks lag at 45 to 55 percent,
and a model’s performance on common conditions poorly predicts its performance on rare ones — the
high-consequence cases clinicians are trained never to miss. Licensing also relies on something a
model lacks: a physician who is unsure consults a colleague, checks a reference, or refers. The device
counterpart is deferral to human interpretation, which, unlike competency, can be tested directly.
Take the comparator from the standard of care where the device will be used. For a function used by
a specialist, it is the specialist with versus without the device. For an autonomous function, it is a panel
of qualified specialists where specialists are available, and the median specialist in practice where they
are not — never the absence of interpretation, because an unavailable alternative changes the benefit
of a device, not the consequence of its errors. Thresholds should be prespecified and absolute, as for
the first FDA-authorized autonomous diagnostic device, and the same whether or not a device is
generative.
C. Postmarket monitoring (Section V; Questions 18–20, 24, 25)
Treat postmarket monitoring as a complement to premarket evidence, not a substitute for it.
Reduced premarket evidence does not reduce uncertainty; it transfers its discovery to patients, while
the benefit of earlier market entry accrues to the sponsor. Monitoring could replace premarket evidence
only where failures are detectable in monitored data, caught before they cause harm, automatically
mitigated, and attributable to a device with reproducible output — conditions that rarely hold together
for high-risk functions. Monitoring should scale with the reasons premarket evidence is inevitably
incomplete: model evolution, output variability, input drift, and open-endedness; a device with
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 4 of 45
Docket No. FDA-2026-N-7874 · Contents
deterministic, bounded output is exposed chiefly to drift. Reproducibility must be demonstrated, not
configured: a decoding temperature of zero does not guarantee identical output in production serving.
D. Architectural safeguards for high-risk functions (Section VI; Questions 4, 11, 18, 26)
Require architectural safeguards for high-risk functions, because they make failure bounded and
measurable. Accuracy can be measured for any function. What cannot be done for a high-risk function
is to enumerate the inputs on which it will fail — and a control cannot be written against a failure that
cannot be described. We propose three safeguards:
• Quantified uncertainty at each processing stage, validated against observed error, so that an
unstable conclusion is recognized before it is reported.
• An independent secondary evaluation path (safety-net), independent of the primary path in task,
training data, and architecture, to catch the confident error that uncertainty alone cannot detect.
• Structured deferral to human interpretation when a validity check fails, a finding falls outside the
authorized set, the two paths disagree, or uncertainty exceeds a calibrated threshold — so that
what a function cannot characterize is deferred rather than silently omitted.
Performance can then be specified as the error rate among finalized cases, by clinical consequence,
alongside the deferral rate. The safeguards should scale with the risk profile, from conventional
validation alone for lower-risk functions to all three for the highest, whether supervised or autonomous.
E. Beyond generative AI (Section VII; Questions 1, 17, 26)
Apply one risk framework to all AI-enabled device functions, and let substantial equivalence turn on
characteristics, not labels. Many of the paper’s questions, and most of our proposals, are not specific
to generative AI. Several also remain unaddressed by FDA guidance for conventional AI-enabled
devices — how the user bears on risk, for example, and how to evaluate functions whose output is not
reviewed by the clinician who would ordinarily perform the task. Autonomous functions are the clearest
case: the first that FDA authorized was not generative. One framework does not mean one risk profile:
for the same intended use, a function with deterministic, bounded output presents a substantially lower
risk profile than one with stochastic, open-ended output. Classification should still follow intended use,
but a generative implementation with non-reproducible or open-ended output, relative to a
reproducible, bounded predicate, raises a different question of safety and effectiveness that the
predicate’s data cannot answer, and in our view should not be found substantially equivalent to it.
F. Generalizability and independent benchmarking (Section VIII; Questions 10, 12, 13, 16)
Measure generalizability directly, with validated synthetic benchmarks. As devices grow more
accurate on average, what matters is where their remaining errors fall. Real-world test sets cannot
show this reliably: five hundred cases spread across eighteen combinations of equipment, protocol, and
patient age leave fewer than thirty in each — too few to measure performance with confidence, for
common presentations as well as rare ones — and more real-world or non-US data enlarges a test set
without filling its gaps. Synthetic data complements real-world data by filling them. Cases can be
generated for every combination the intended use specifies, varying one factor at a time against a
reference standard known exactly; a fresh test set for each evaluation cannot have been seen in
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 5 of 45
Docket No. FDA-2026-N-7874 · Contents
training; and testing can be automated and repeated before submission, after each change, and after
marketing.
Enable independent third parties to provide these benchmarks, and base confidence on how an
evaluation is run. Third parties bring capability many sponsors cannot build in-house, independence
between development and testing, and a common yardstick across devices. An evaluation that is prespecified, automatically scored, reported identically to FDA and the sponsor, and auditable by FDA is
verifiable by every party, whoever runs it. We recommend that FDA recognize validated synthetic
benchmarks as evidence of generalizability and set MDDT qualification expectations for them.
A note on evidence. Where the literature contradicts or complicates a position we take, we have
cited it and said so. Where a claim rests on reasoned argument rather than direct empirical support,
we have identified it as such, and we have identified preprint and non-peer-reviewed sources. We
take this approach because a comment that overstates its support is of little use to the Agency.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 6 of 45
Docket No. FDA-2026-N-7874 · Contents
III. Comments on Section IV — Considerations for the
Assessment of Risk
RESPONSIVE TO QUESTIONS 1 THROUGH 6.
We support CDRH’s risk-based approach to GenAI-enabled device functions, and in particular its
intention to distinguish among informational functions rather than treating informational output as
uniformly low in risk. We offer a proposed refinement. We believe it builds directly on considerations
the discussion paper itself raises, on the traceability dimension that Question 1 invites, and on the risk
categorization framework of the International Medical Device Regulators Forum, which FDA adopted in
its 2017 guidance on the clinical evaluation of software as a medical device. 2,3
A. Considerations the discussion paper already identi2es
In several places the discussion paper qualifies its two-axis structure in ways we find persuasive. We
set them out here because our proposal is, in large part, an attempt to give them a structural place.
On the directiveness of informational outputs based on the substance and context of the output.
CDRH is considering “whether the degree to which an informational function may be directive may
depend on the substance and context of the output, not solely on whether it uses words such as
‘recommend,’ ‘should,’ or ‘consider,’” and whether a patient-facing function “may not become any
less directive because it includes a ‘talk to your doctor’ or an ‘I am not a medical professional’
statement.”
On output basis evaluability as a risk factor. CDRH observes, of measurement and signalprocessing functions, that “although these functions yield non-directive information, they may
nonetheless be of higher risk because the user typically cannot independently evaluate the basis
for the output (and therefore the function may migrate higher on the consequences axis).”
On the user. CDRH notes that health care professionals “may be better equipped than patients to
interpret an output in context, to recognize its limitations, and to integrate it with other clinical
information such that they can identify an incorrect output and not rely upon it,” and that “a
function whose safe use depends on contextualization by specialist expertise may present
elevated risk when used by HCPs lacking that expertise to independently evaluate outputs.”
We would add one related observation, concerning the paper’s own example of non-directive
information: “a risk score for a future cardiovascular event.” The same function, returning a very high
estimated risk for a particular patient, will predictably prompt further evaluation and intervention. The
function has not changed; the clinical consequence of its output has. A function’s position on the
activity axis may therefore depend on the output actually delivered, and not only on the function’s
description.
B. A common source
Several of these considerations appear to share a single origin. The discussion paper organizes its first
axis around the degree to which a function directs or takes action, and notes that the framework draws
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 7 of 45
Docket No. FDA-2026-N-7874 · Contents
conceptual structure from analogous matrices in other FDA guidance, including guidance on the
credibility of computational modeling and simulation and draft guidance on the use of artificial
intelligence in drug and biological product regulation.4,5 Those matrices organize their corresponding
axis around model influence — the contribution of a model’s output to the decision at hand. The IMDRF
framework organizes it around the significance of the information provided to the healthcare decision. 2
Both are questions of significance. The discussion paper’s activity axis recasts them as a question of
directiveness.
The difference is consequential. Directiveness is naturally read as a property of how an output is
expressed, which is why the question of wording arises. A measurement may be non-directive in
expression yet highly significant to a diagnostic decision, which is why measurement functions sit
uneasily on the axis. And directiveness may vary with an output’s value, as the risk-score example
illustrates. Significance, by contrast, attaches to what an output asserts. It is stable across phrasing, and
more stable across output values.
The considerations concerning the user are different in kind. Neither directiveness nor significance
describes who receives an output, or what that person is equipped to do with it. We address the user
separately below, as a distinct dimension.
C. A proposed re2nement: four dimensions of risk
We propose that the risk of a device software function be assessed along four dimensions. The first
refines the activity axis. The second is retained. The third gives structural place to a factor the
discussion paper identifies but has not yet positioned. The fourth makes explicit the technological
characteristics that the paper itself identifies as the source of the evaluation challenges GenAI presents.
Dimension The question it asks Known from
1. Significance of the Does the function present or generate Design and intended
output clinical-decision content (clinical claims), use
and if it generates, what kind of claim?
2. State of the condition How severe is the condition to which the Intended use
output pertains?
3. The user of the output Can the recipient evaluate the output and Intended user and
its basis, and act on it correctly? design
4. Characterizability Is the output reproducible, and are the Design
input and output spaces bounded?
1. Significance of the output
Question 1 asks whether additional dimensions should be represented in the framework, and names
among them “the traceability of the output (i.e., to primary source materials).” We believe traceability is
not merely an additional dimension but the most useful organizing principle for the activity axis itself,
and we propose a formalization of it: an axis defined by the significance of the output, organized
around whether a function presents clinical-decision content (clinical claims), or generates it.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 8 of 45
Docket No. FDA-2026-N-7874 · Contents
A function presents clinical-decision content (clinical claims) when that content originates elsewhere —
in the function’s own input, or in an identified external source — and the function conveys, summarizes,
reformats, or applies it. A function generates clinical-decision content (clinical claims) when the content
originates with the model: the function produces a clinical claim concerning detection, diagnosis, or
clinical management that did not exist as a claim in any identified source.
This distinction corresponds closely to the boundary IMDRF draws between informing and driving
clinical management. IMDRF describes informing clinical management as providing information that
“will not trigger an immediate or near term action,” including “to provide clinical information by
aggregating relevant information.” It places functions intended to aid in diagnosis, to aid in treatment,
or to triage or identify early signs of a disease or condition in the category of driving clinical
management, and functions that diagnose, screen, or detect in its highest category. 2 On IMDRF’s own
terms, a function that contributes a clinical claim of its own — even as an aid — is no longer informing.
We therefore propose that presenting functions constitute the informational category, and that
generating functions be graded according to the kind of claim they generate:
Level Significance of the output Category
1 Presenting clinical-decision content (clinical claims) Informational
2 Generating a quantitative measurement Generating
3 Generating a detection Generating
4 Generating a diagnosis Generating
5 Generating a clinical-management determination, Generating
including treatment or treatment adjustment
The ordering reflects the increasing directness with which a generated claim determines clinical action.
We note that IMDRF places detection, diagnosis, and treatment within a single highest category. The
finer grading proposed here is intended to reflect differences in the evidence each warrants, and we
would welcome CDRH’s view on whether the ordering among the upper levels should be strict.
This structure also clarifies one feature of the IMDRF axis. That axis combines two considerations: the
type of clinical determination involved, and whether an output is offered as an aid to a clinician’s
determination or as the determination itself. We propose separating them. The significance axis grades
the type of claim a function generates. Whether a qualified person evaluates that claim before it has
effect is a separate matter, which we address under the third dimension.
Three clarifications make the distinction operable.
Raw data is not a claim. Images, physiological signals, laboratory values, and descriptions of symptoms
are inputs to clinical judgment, not clinical judgments. A function that produces a diagnosis from an
image has generated that diagnosis, however much information the image may be said to contain.
A clinical claim has three components. What makes an assertion a clinical claim is its subject matter:
detection, diagnosis, or clinical management. The three components identify where, within such a claim,
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 9 of 45
Docket No. FDA-2026-N-7874 · Contents
the model’s own judgment may enter. A function presents only where none of the clinically decisive
components originates with the model.
• Its content, including the certainty with which it is asserted. A function that restates a
radiologist’s “cannot exclude early ischemia” as “early ischemia is present” has generated the
definite claim, although the entity itself was drawn from the source.
• Its applicability to the particular patient. The Fleischner Society guidelines for incidental
pulmonary nodules state their own scope, excluding patients younger than 35 years, patients
with known malignancy, immunocompromised patients, and lung cancer screening. 6 A function
that conveys their recommendation accurately, but for a 28-year-old or for a patient with known
melanoma, has generated the determination that the guideline applies — and generated it
wrongly — although every word of the recommendation is correct. The separation of applicability
from content mirrors a long-established distinction in evidence-based medicine, which treats
whether evidence applies to the patient at hand as a question distinct from whether the evidence
is valid and what it shows.7
• Where several findings bear on one conclusion, their synthesis. Synthesis is presenting only
where the source itself specifies how the findings combine, as diagnostic criteria and scoring
systems do. Where it does not, as in most diagnostic reasoning, the synthesis is generated. A
function that combines a dilated appendix, periappendiceal fat stranding, and an appendicolith
into a diagnosis of appendicitis has generated that diagnosis, even if each finding was drawn
from the radiologist’s own description.
The distinction applies per claim, not per device. A single output may contain both: for example, a
measurement the function generated, stated alongside the published reference range for that
measurement. This is consistent with CDRH’s approach to multiple function device products, 8 and it
places the evidentiary burden precisely where the clinical judgment was made.
The underlying construct is not novel to us. The distinction between output verifiable against an
identified source and output originating with the model has been formalized in the natural language
generation literature as attribution to identified sources, with an operational test asking whether a
statement can be affirmed in the form “According to [source], [statement].” 9 The related distinction
between faithfulness to a source and factual correctness in the world,10 and the distinction between
intrinsic content, which contradicts its source, and extrinsic content, which its source neither supports
nor contradicts,11 are both well established. Extrinsic content corresponds closely to what we call
generated content. It may be factually correct; what makes it generated is its origin, not its accuracy.
We draw on this work because it supplies validated methodology for the question CDRH now faces.
We wish to emphasize one point, because it is where the distinction is most likely to be misapplied. A
citation attached to an output does not make the output presenting. A generative model can produce
a statement naming a guideline, a reference, and a figure, where the statement itself was produced
from the model’s learned parameters rather than drawn from the source it names. An empirical
evaluation of retrieval-grounded professional research tools marketed as eliminating unsupported
content found hallucination rates of 17 to 33 percent,12 and benchmarking of citation generation has
found that leading systems frequently lack complete support for the statements they cite. 13 What makes
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 10 of 45
Docket No. FDA-2026-N-7874 · Contents
a function presenting is that its clinical claims are demonstrably drawn from the identified source. How
a sponsor establishes this is a matter of design; that it can be established is the requirement.
Finally, we wish to be precise about what this distinction does not do. It does not determine whether a
software function is a device; that question is governed by section 201(h) of the FD&C Act and the
exclusions in section 520(o)(1). We note, however, that the distinction tracks the boundary drawn by
section 520(o)(1)(E) closely, as the paper’s footnote 13 suggests for the related consideration of
measurement functions. A presenting function directed to a health care professional who can
independently review the basis for its output may, depending on its other characteristics, fall within that
provision; a generating function generally cannot, because the basis for its claim is the model’s own
inference.14 Within the space of device functions, we propose that a function’s position on the
significance axis inform the evidence expected of it.
2. State of the healthcare situation or condition
We propose retaining the consequence axis. The severity of the condition to which an output pertains is
an essential determinant of risk, well established in the IMDRF framework as the state of the healthcare
situation or condition — critical, serious, or non-serious.2 A generated diagnosis concerning a selflimiting condition and one concerning an acute cerebrovascular event do not carry comparable risk, and
the significance axis alone does not capture that difference. Two of the further considerations Question
1 names — the reversibility of a resulting action and the time pressure of the deployment setting — are,
in our view, naturally expressed on this dimension, since both bear on the severity of harm that follows
from relying on an incorrect output.
3. The user of the output
The discussion paper identifies the user as relevant to risk but does not yet give the user a structural
place. We propose that the user be assessed through three considerations. Two concern the user. The
third concerns the device, but determines how much the user is able to contribute to safe use.
Evaluation of the output. Can the user determine whether the output itself is correct? For subspecialty
content, generally only the corresponding specialist can. The evidence that this matters is strong. In a
prospective study of radiologists reading mammograms with purported artificial intelligence assistance,
incorrect suggestions reduced correct assessment among inexperienced readers from approximately
80 percent to approximately 20 percent, and less-experienced readers were significantly more affected
than experienced ones.15
Evaluation of the basis for the output. Can the user understand how the output was produced well
enough to judge independently whether it may be wrong? This is the consideration the discussion paper
raises in connection with measurement functions, and the consideration on which section 520(o)(1)(E)
turns. It is determined principally by the design of the device rather than by the user. We would add that
it must be demonstrated rather than assumed. In a randomized study of 457 clinicians, systematically
biased model predictions reduced diagnostic accuracy by 11.3 percentage points, and the addition of
image-based explanations improved accuracy by only 2.3 points.16 We therefore suggest that where a
sponsor offers explainability as a basis for reduced risk, it be benchmarked against its purpose:
whether it measurably improves the intended user’s ability to detect incorrect outputs.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 11 of 45
Docket No. FDA-2026-N-7874 · Contents
Applicability. Can the user act correctly on the output within the established clinical pathway? This
depends on whether the user is the person who would ordinarily receive and act on that category of
information. A subspecialty imaging finding delivered to the clinician who ordered the examination
enters a pathway in which that clinician routinely contextualizes and acts on such findings. The same
finding delivered to a clinician outside that pathway, or to a patient, does not — even though the output
is identical.
The discussion paper suggests that patient-facing considerations could be captured by shifting certain
functions higher on the consequences axis. We would respectfully suggest a dedicated dimension
instead. The user does not change the severity of harm that follows from an incorrect output; it changes
the likelihood that an incorrect output will be recognized and not acted upon. These are distinct
mechanisms, and representing the second as the first would require the consequences axis to carry
two meanings.
We share the discussion paper’s recognition that functions directed to patients and to generalist
clinicians can bring real benefits, including extending access to specialty knowledge and supporting
patient engagement, and we do not regard these recipients as inherently inappropriate. The purpose of
this dimension is not to discourage such functions, but to identify what they must demonstrate.
We believe this dimension also accounts for the distinction between supervised and autonomous
operation, so that a separate position for autonomy on the activity axis is not required. In the sense that
matters for risk, a function operates autonomously when the person best placed to evaluate its output is
not the person who receives it, or when no one receives it before it takes effect. That is a statement
about the three considerations above. A fully automated subspecialty output delivered to the
corresponding specialist, to a generalist, and to a patient is equally automated in each case, yet the risk
differs substantially among them, and these three considerations describe why. Where no qualified
person receives an output before it takes effect, all three are at their maximum by construction. The
availability of downstream safeguards, which Question 1 also names, is expressed here: it is the
question of whether a qualified person evaluates the output before it has effect.
4. Technological characteristics bearing on characterizability
We propose a fourth dimension comprising two properties of a function’s design: the determinism of its
output, and the bounding of its input and output space. We group them because they share a single
rationale. Each determines how completely premarket testing can characterize the function’s behavior,
and therefore how much evidence is needed to do so.
Output determinism. A function may be specified to return identical output for identical input, or its
output may vary across runs. This can be characterized from the design before any testing is
performed: whether output is specified to be reproducible, and by what mechanism — deterministic
execution in a configuration fixed for that purpose, generative inference with controlled decoding, or
generation by sampling.
Bounding of the input and output space. A function’s inputs may be constrained to a defined
specification, with out-of-specification inputs rejected, or it may accept open-ended input. Its outputs
may be constrained to a defined set of claims or a defined structure, or it may produce open-ended
output. This, too, is a property of the design.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 12 of 45
Docket No. FDA-2026-N-7874 · Contents
A dimension of risk should be recognizable from the design, because risk determines the evidence that
will be required, and the evidentiary burden cannot depend on the results of the evidence. For this
reason we distinguish this dimension from the reproducibility of a function’s clinical claims, which is an
important endpoint of validation but is known only once validation has been performed. The two stand
in a clear relationship. Greater output variability, and a larger admitted input and output space, raise the
evidentiary burden: the sponsor must show that output variability does not translate into variability of
clinical claims, and that performance holds across the space the function admits. Claim-level
reproducibility, measured as described in Section IV.D, is how that burden is discharged. A highly
stochastic function may, on evaluation, prove to make reproducible clinical claims; it is precisely
because this must be shown that the dimension belongs in the assessment of risk.
This dimension interacts with the others, as the others interact with one another. Variability matters in
proportion to what varies. Variable phrasing in a presented summary is of limited consequence; a
generated diagnosis that varies across runs is a serious failure, as the discussion paper recognizes in
stating that “variation in safety-critical behaviors such as escalation, refusal, or diagnostic conclusions
is treated as a failure.”
D. Faithfulness as a benchmarking element for presenting functions
RESPONSIVE TO QUESTION 9.
A presenting characterization should not be self-certifying. Whether a function presents is a question
about the origin of its claims, and can be determined from its design and intended use. Whether it
conveys its source accurately is an empirical question. We recommend that CDRH treat faithfulness —
accurate conveyance of the identified source, without addition, alteration, omission of material
conditions, or unsupported extrapolation — as a required benchmarking element for presenting
functions. Validated methodology exists for this purpose, including the attribution annotation framework
developed in the natural language generation literature and entailment-based measures of faithfulness
developed in the summarization literature.9,10
E. Multi-turn operation and escalation
Question 5. If the relevant distinction is whether a function generates clinical-decision content (clinical
claims), migration over the course of a conversation is resolved directly. A function that generates such
content at any point in an exchange occupies the corresponding position on the significance axis for
that exchange, regardless of how many turns preceded it, and its intended use can be characterized
accordingly.
Question 6. We suggest that under-escalation and over-escalation need not be treated as
incommensurable. Where a function generates an escalation recommendation for a defined clinical
condition, correct and incorrect escalation constitute a classification problem, and the two directions of
error correspond to sensitivity and specificity for that condition. Sponsors can reasonably be expected
to prespecify both, for each condition within the indication for use, and to justify the relative weight
given to each by reference to the clinical consequence of each type of error.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 13 of 45
Docket No. FDA-2026-N-7874 · Contents
F. Scope
Several of the considerations above — the position of measurement functions, and the dependence of
directiveness on output value in particular — arise for conventional device software functions as well.
Generative functions make them more visible, because text output lends itself to reading directiveness
as a matter of wording. We believe the refinement proposed here would improve the consistency of the
framework generally, and we would encourage CDRH to consider it with that broader scope in mind.
We wish to be clear, however, that a common framework does not imply a common risk profile. For the
same intended use, a conventional function with fixed weights, output configured to be reproducible,
and bounded inputs and outputs sits low on the fourth dimension, while a generative function with
variable output and open-ended inputs and outputs sits high on it. The conventional function will
therefore present a substantially lower risk profile, and the generative function will warrant substantially
more evidence, even where the two share an intended use.
Illustrative examples of the generating and presenting distinction appear in Appendix A, and of how the
four dimensions combine in Appendix B.
IV. Comments on Section V — A Competency-Based
Approach for Premarket Evaluation
RESPONSIVE TO QUESTIONS 7 THROUGH 17.
We welcome CDRH’s statement that a competency-based approach “could be tailored to the device’s
intended use and be proportionate to its risk.” Our comments concern the premise on which the
approach rests, the limits of inferring competency, an alternative that we believe makes much of the
approach unnecessary for many devices, the proportionality of evidence to risk, and the selection of
comparators.
A. A conditional premise
RESPONSIVE TO QUESTIONS 7 AND 9.
The methodological innovations proposed in Sections V and VI respond, in large part, to two properties
the discussion paper attributes to GenAI-enabled devices: the open-endedness of their inputs and
outputs, and the variability of their outputs in response to similar inputs. These properties recur as the
stated rationale throughout.
The competency-based approach. The paper observes that “evaluation approaches developed
for software with bounded inputs and fixed outputs may not be appropriate for GenAI-enabled
devices,” because “the range of possible inputs and outputs may be too large for such testing to
be practical”; that “a single correct response often does not exist and many different responses
may be acceptable”; and that clinicians and GenAI-enabled devices alike “can ingest broad and
varied information and can produce open-ended responses across a wide range of situations.”
Adjudication by language models. Evaluation of open-ended output at scale requires adjudication
at scale, and the paper contemplates that adjudication may be performed by language models,
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 14 of 45
Docket No. FDA-2026-N-7874 · Contents
noting that its independence considerations “would still be applicable when the expert adjudicator
is itself an LLM.” We support the extension of independence requirements to such adjudicators.
Postmarket monitoring. The paper states that these devices “produce varied, open-ended
outputs in response to similar inputs and may undergo continuous adjustments following
deployment, making it difficult for premarket testing to fully capture performance.”
Elsewhere, the paper frames these properties conditionally. Section II states that GenAI-enabled
devices “may accept open-ended inputs” and “produce variable outputs to similar inputs,” and
describes the premarket challenge as arising for “devices with broad intended uses or those that use
open-ended input and output formats.” We believe this conditional framing is correct, and important.
Open-endedness and output variability are properties a device may have; to a considerable extent, they
are properties a device is designed to have. A device that constrains its inputs, constrains its outputs,
and controls the variability of its output does not present the evaluation problem the competency-based
approach is designed to address.
We therefore recommend that the framework ask sponsors to establish whether these conditions hold
for their device, rather than presume that they do. A device that has bounded its input and output space
and controlled its output variability should be evaluable by established methods. A device that has not
should bear the correspondingly greater evidentiary burden. This allocates the burden to the party that
controls the design, and it rewards the design choices that make devices most evaluable.
B. The limits of competency inference
RESPONSIVE TO QUESTIONS 7 AND 9.
The competency-based approach is modeled on medical licensing, board certification, and supervised
practice. We think the analogy warrants careful examination, because the available evidence indicates
that performance on licensing-style assessment does not presently transfer to open clinical practice for
these systems.
A systematic review of 39 medical language model benchmarks found licensing-examination-style
benchmarks approaching saturation, with leading models scoring in the mid-eighties to low nineties,
while practice-oriented assessments lagged substantially — diagnostic tasks in the range of 45 to 55
percent and safety assessment in the range of 40 to 50 percent.17 Related work documents accuracy
degradation when clinically irrelevant information is added to otherwise identical cases. 18 Benchmark
contamination is a recognized and material concern; one clinical evaluation team withheld the majority
of its dataset from publication specifically to guard against it.19
We do not offer this as an argument that structured assessment is without value. We offer it as
evidence that inferring open-practice performance from structured-assessment performance is not
currently supported for these systems, and that a framework resting on that inference should state what
would make the inference reliable.
There is a second, more specific concern. Physician credentialing functions in part because clinical
training weights low-prevalence, high-consequence diagnoses independently of their base rate; the
discipline of the “must not miss” differential is taught precisely because unaided statistical reasoning
would underweight it. Model behavior on rare presentations is instead determined by their
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 15 of 45
Docket No. FDA-2026-N-7874 · Contents
representation in training data, and performance on common conditions is a poor predictor of
performance on rare ones — a pattern documented in long-tailed medical imaging, where methods that
improve rare-class performance typically do so at a cost to common-class performance. 20 Competency
inference is therefore least reliable precisely where clinical consequence is greatest. We are not aware
of a direct empirical comparison of must-not-miss clinical reasoning with base-rate-driven model
behavior, and we offer this as reasoned argument rather than as a settled finding.
Physician credentialing also presumes a second property: a physician who does not know can consult,
refer, or defer, and licensure is granted on the expectation of that behavior. We anticipate the response
that agentic systems partially reproduce this through retrieval and tool use. The transferable analogue,
however, is deferral — and deferral is directly testable in a way that inferred competency is not. We
regard this as an argument for making deferral behavior a validated device property rather than for
accepting competency inference.
C. Bounding the input and output space
RESPONSIVE TO QUESTIONS 7, 9 AND 17.
The point is easily illustrated with conventional artificial intelligence. A network trained on a specific
imaging modality will accept an arbitrary input array and produce some output; its theoretical input
space is unbounded. We do not conclude from this that such a device must be evaluated by inference
from adjacent competencies. We require the sponsor to define the input space, to reject inputs outside
it, and to test within it. The unbounded theoretical capability of the underlying model has never been
treated as a reason to relax evaluation of the bounded device.
Comparable instruments exist for generative architectures. Input may be bounded by restricting a
function to a pre-specified set of prompts or task definitions rather than free-form instruction; by
constraining the sources and formats of input; by rejecting out-of-specification input; and by restricting
retrieval to a curated and versioned corpus. Output may be bounded by restricting clinical claims to
selection from a pre-specified set; by generating within a structured report template; by constraining
decoding to a defined schema, a technique demonstrated across structured tasks and implemented in
widely used inference frameworks;21,22,23 and by screening outputs with secondary validation models.24,25
Output variability may be reduced through the configuration of decoding, subject to the technical
caveat recorded in Section V.C. Worked illustrations appear in Appendix C.
These instruments suit some intended uses better than others. A pre-specified set of prompts costs a
function that drafts imaging reports very little, because its principal input is the image; it would remove
most of the value of a function whose purpose is open conversation. We regard this as the point rather
than an objection to it. Where a sponsor can bound a function, it should be expected to, and should
then bear a correspondingly lighter evidentiary burden. Where a sponsor cannot, or chooses not to, the
function sits higher on the fourth dimension of risk proposed in Section III, and the greater burden
follows.
We think it important to state the limitations of these methods rather than overstate their sufficiency,
because the limitations bear directly on how CDRH should treat a bounding claim.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 16 of 45
Docket No. FDA-2026-N-7874 · Contents
1. Constraint guarantees structure, not clinical correctness. A grammar can compel an output into
a valid form; it cannot make the value correct.
2. Constraint can degrade output quality. Token-level masking has been shown to distort a
model’s output distribution, yielding outputs that are structurally valid but not representative of
the model’s own likelihoods, and benchmarking against real-world schemas shows substantial
residual failure.26,27
3. Constraint mechanisms are themselves an attack surface. Recent work demonstrates that
structured-output interfaces can be exploited to circumvent model safety behavior. 28
The conclusion we draw is not that bounding is unreliable. It is that bounding mechanisms are device
design features requiring their own verification and validation, and that a bounding claim should be
evidenced rather than asserted. Bounding does not establish correctness. It makes rigorous evaluation
tractable and the failure space enumerable, which is what a regulatory framework requires in order to
specify controls at all.
We therefore recommend that CDRH treat bounding of the input and output space as the primary
regulatory expectation, and require sponsors proposing competency-based evaluation to justify why
bounding is not achievable for the function at issue. The discussion paper already moves in this
direction where it states that benchmarking should be “of sufficient clinical and functional scope to
evaluate the device for its intended use,” extending where appropriate to clinical domains beyond those
a device nominally addresses. We read that as an acknowledgment that scope must be controlled
rather than assumed.
We also wish to bound our own argument. Exhaustive testing is feasible only where reference data
exist, and for genuinely rare conditions they often do not. Our position is that where a condition cannot
be evidenced at the required level, the appropriate response is to exclude it from the authorized output
set and route it to human interpretation — not to infer acceptable performance from competence in
adjacent domains.
D. Translating open-ended output into clinical claims
RESPONSIVE TO QUESTIONS 9, 10, 12 AND 14.
Our central recommendation for the premarket evaluation of generating functions is this: GenAIenabled devices do not require a new category of evidence. They require that open-ended output be
translated into clinical claims that established forms of valid scientific evidence can evaluate.
FDA’s regulations define valid scientific evidence to include well-controlled investigations from which
qualified experts can fairly conclude that there is reasonable assurance of safety and effectiveness. 29
For a diagnostic function, the most established form of such evidence measures the accuracy of the
function’s determinations against a reference standard. That form of evidence does not depend on
whether an output is phrased as free text. It depends on whether the clinical claims the output makes
can be identified. A report that states a diagnosis has made a claim about that diagnosis; the claim is
correct or incorrect against the reference standard, and its sensitivity and specificity can be measured
and bounded with confidence intervals by conventional methods. The discussion paper acknowledges
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 17 of 45
Docket No. FDA-2026-N-7874 · Contents
that the approaches it describes “may not be powered around traditional effectiveness endpoints.”
Claim-level evaluation is powered around exactly those endpoints.
This connects directly to our comments on Section IV. The distinction between generating and
presenting identifies which content in an output is a clinical claim originating with the device. Evaluation
then asks two questions of each such claim: whether it is correct, and whether it is reproducible.
Open-endedness: mapping free text to a closed space of clinical claims
Open-ended output can be mapped to a closed space of clinical claims in two ways, which differ in an
important respect.
Structure first. The generating function emits its clinical claims in structured form as its primary output
— by constrained decoding into a defined claim schema, or by selection from an allow-listed set of
claims — and any free text is rendered from that structure. The structured claims are then authoritative,
and are evaluated directly against the reference standard. No extraction step is required, and no
language model is needed to judge what the output asserted.
Text first. The generating function emits free text, and a claim extractor identifies the clinical claims it
contains. The extractor is then itself a component of the evaluation, and its performance must be
validated against human extraction, for both accuracy and coverage. This is so whether the extractor is
rule-based or model-based: a rule-based extractor is reproducible but may miss claims expressed in
unanticipated phrasing, so its coverage must still be shown.
We recommend the structure-first approach as the preferred mechanism for functions at the higher end
of the risk profile. It removes a source of measurement error, it makes the claims a function has made
unambiguous, and it reduces reliance on language-model adjudication. The text-first approach remains
appropriate where a sponsor can show that its extractor is adequately accurate and complete.
Variability: measuring reproducibility at the level of the claim
Each input is evaluated repeatedly under the deployed configuration, and the clinical claims are
identified in each output. Agreement is measured at the level of the claim: for each finding, the
proportion of runs in which the function makes the same claim. The same measurement is repeated
under clinically irrelevant perturbation of the input — paraphrase, reordering of information, the addition
of irrelevant detail — which the discussion paper identifies among the considerations for element R.1.
This is the endpoint that discharges the evidentiary burden set by the fourth dimension of risk in Section
III: output determinism is known from the design and determines how much must be shown; claim-level
reproducibility is what shows it.
Three limitations
The extractor is a model. Where the text-first approach is used, extraction relocates rather than
removes the need for automated judgment, which is why its validation is required.
Claims outside the defined set escape evaluation. Claim-level evaluation assesses the claims it was
designed to identify. An output that makes an unanticipated clinical claim would not be scored. A
complementary check is therefore needed: every clinical claim in an output must map to the authorized
claim set or be flagged, which corresponds to element S.2, scope maintenance and boundary
adherence.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 18 of 45
Docket No. FDA-2026-N-7874 · Contents
Certainty is part of the claim. “Possible” and “definite” findings of the same entity are different claims,
and extraction must capture the qualification as well as the entity. Omission is handled naturally: a
finding the output fails to report is a false negative against the reference standard.
Question 10. The discussion paper notes that sponsors might “propose their own specialized tests and
datasets tailored to their specific device and its functions,” and Question 10 asks what role such
sponsor-developed benchmarks should play given concerns about independence and optimization to
the test. Translating output into clinical claims allows those claims to be evaluated against established,
independently adjudicated reference standards, which reduces reliance on sponsor-developed
benchmarks whose construct validity would otherwise have to be separately established.
Question 12. Because claim-level endpoints are conventional diagnostic accuracy endpoints, sample
sizes and confidence intervals can be determined by conventional methods, per claim and per stratum
of clinical consequence.
E. Evidence proportionate to risk pro2le, and the role of architecture
RESPONSIVE TO QUESTIONS 8, 11 AND 17.
Question 8 asks how the risk framework might inform the level of evidence required. We recommend
that the four-dimension risk profile described in Section III determine the nature, rigor, and amount of
evidence: which benchmarking elements apply, the range and difficulty of test conditions, the extent of
clinical confirmation, and, as discussed in Section VI, which architectural safeguards are expected. The
discussion paper already anticipates this, noting that benchmarking elements “would be chosen based
on applicability to the device’s intended use and risk profile.”
Architecture determines which benchmarking approaches are applicable. We would suggest the
following correspondence:
Characteristic of the function Benchmarking approach it calls for
Generates clinical claims Claim-level accuracy against a reference standard
(Section IV.D)
Presents clinical claims Faithfulness to the identified source (Section III.D)
Variable output Claim-level reproducibility across runs and under
perturbation (element R.1)
Open-ended input or output Scope maintenance and detection of claims outside the
authorized set (element S.2)
Conversational interface Multi-turn, escalation, communication, and paraphrase
elements (S.1, E.2, E.4, R.1)
Image input; measurement or Established quantitative imaging methods, including
segmentation output repeatability and reproducibility
This bears directly on Question 17, which asks whether aspects of the approach might be ineffective or
inapplicable to other model architectures. For imaging devices, much of Appendix A presupposes a
conversational interface. Multi-turn simulated encounters, the sequencing of follow-up questions,
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 19 of 45
Docket No. FDA-2026-N-7874 · Contents
resistance to over-reassurance when users minimize symptoms, adversarial and emotionalmanipulation prompting, empathy and therapeutic appropriateness, health literacy, consistency across
semantically equivalent paraphrases, and performance across dialects and accents are all meaningful
for a device that converses with a user. A function that accepts a volumetric image series and returns a
measurement or a segmentation has no conversational surface on which they can be assessed.
Applying them would impose burden without safety return, contrary to least burdensome principles.
We would add one observation on the function taxonomy in Section IV of the paper. Measurement
functions and diagnosis are each addressed. Detection is not addressed as a function type. In imaging,
detection is where autonomy risk is concentrated, because a missed finding produces no error signal:
the output is confident, complete in appearance, and contains no indication that anything was omitted.
We recommend that detection be addressed explicitly.
Where an existing classification regulation and its associated guidance already govern a function, a
generative implementation should not displace them. A function returning a quantitative imaging
measurement for clinician interpretation should remain subject to the requirements applicable to that
device type, including the repeatability and reproducibility assessment that existing quantitative imaging
guidance already describes.30 That guidance addresses output variability directly, using established
methodology, without the need for a parallel framework.
F. Comparator selection and acceptance criteria
RESPONSIVE TO QUESTIONS 14 AND 15.
The discussion paper frames comparator selection with care. It suggests that performance might be
compared “to that of a panel of qualified clinicians whose consensus reflects the applicable standard
of care, or to that of a median clinician in practice”; that “considerations for selecting the comparator
might include how the device is actually used”; and that benchmark datasets be “representative of the
deployment contexts in which the device is expected to be used.” We agree that no single comparator
suits all devices, and that the choice should depend on the intended use and the deployment
environment. The paper identifies the relevant alternatives: a panel of qualified clinicians, a median
clinician in practice, generalist or specialist physicians, and the performance of a human-AI team as
against the device working alone.
We propose one organizing principle: the comparator should be the standard of care in the environment
for which the device is indicated. Applied to device software functions that bear on specialist
interpretation, it yields three comparators.
Intended use and deployment environment Comparator
Function intended for use by a specialist Specialist with the device, compared with
specialist without the device
Autonomous function, where specialist Panel of qualified specialists, compared with
interpretation is available the device operating alone
Autonomous function, where specialist Median specialist in practice, compared with
interpretation is not available the device operating alone
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 20 of 45
Docket No. FDA-2026-N-7874 · Contents
The first is the established reader-study paradigm, and we see no reason to depart from it. In the
second, the prevailing standard of care already includes specialist interpretation, consultation among
specialists, and existing assistive tools. A device that operates in place of that process should be
measured against the consensus of qualified specialists, which reflects it.
The third requires more discussion. Where no specialist is available, the realistic alternative may be no
interpretation, delayed interpretation, or interpretation by a clinician without subspecialty training. It
does not follow that the device should be compared with that alternative. A patient harmed by an
incorrect output is not harmed less because no specialist was available: the absence of an alternative
changes the benefit of the device, not the consequence of its errors. At the same time, a consensus
panel represents a standard of care the setting never offered. We therefore consider the median
specialist in practice the appropriate comparator. It represents the specialist-level care that the setting
lacks and that the device is intended to provide.
We recognize that a median-clinician comparator is a standard that moves with the performance of the
comparison group, whereas an adjudicated reference standard is not. We do not believe that instability
disqualifies it in this setting, but we would recommend two measures to contain it: that the median be
established from a documented and representative sample of practicing specialists, and that it be used
to set a prespecified, absolute acceptance threshold rather than serve as a comparison that moves with
each new study.
This bears on Question 15. The care that would occur in the device’s absence is, in our view, properly
considered in the benefit-risk determination, and we agree with CDRH that it is a legitimate and underused consideration. It should not, however, set the acceptance criterion. The acceptance criterion for
an autonomous function should be an absolute performance threshold, prespecified for each condition
in the indication for use, informed by the comparator above, and justified by the clinical consequence of
a missed or incorrect determination. The comparator establishes where the threshold must at least lie;
the consequence of error may justify a higher one.
This is consistent with precedent. The first FDA authorization of an autonomous diagnostic device,
indicated for use by health care providers to automatically detect more than mild diabetic retinopathy,
rested on prespecified absolute thresholds for sensitivity and specificity of 85.0 percent and 82.5
percent, measured against an independent reading-center reference standard. 31,32
Finally, the choice of comparator should not depend on the implementation technology of the device. A
generative and a conventional device intended for the same use, in the same environment, should be
held to the same comparator and the same acceptance threshold.
V. Comments on Section VI — Postmarket Monitoring
RESPONSIVE TO QUESTIONS 18 THROUGH 20, 24 AND 25.
A. A complement to premarket evidence, not a substitute for it
The discussion paper records CDRH’s earlier view that “it may be important for premarket evidence to
be complemented by postmarket evidence.”33 Section VI asks a different question: whether “it is
appropriate to accept greater premarket uncertainty regarding a GenAI-enabled device’s benefit-risk
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 21 of 45
Docket No. FDA-2026-N-7874 · Contents
profile through greater reliance on postmarket monitoring,”34 and Question 18 asks what monitoring
characteristics would “justify reduced premarket evidence.” The movement from complement to
substitute is significant, and we would encourage CDRH to consider it deliberately.
We respectfully but strongly recommend that postmarket monitoring be treated as a complement to
premarket evidence, and not as a substitute for it. Premarket evidence establishes that a device
performs as intended in a characterized population. Postmarket monitoring establishes that the
conditions under which that finding was obtained continue to hold. These are different questions, and
the second cannot answer the first. Reduced premarket evidence does not reduce uncertainty about a
device’s safety and effectiveness; it transfers the discovery of that uncertainty to the deployed
population, where it is discovered through harm to patients. The benefit of earlier market entry accrues
to the sponsor. The cost of an undetected failure is borne by patients.
For monitoring to compensate for evidence not gathered before marketing, four conditions would need
to hold together:
1. The failure mode is detectable in monitored data. Uncertainty about device behavior on
presentations that do not occur, or occur too rarely to be observed, in the monitored population is
not reduced by monitoring that population.
2. The detection interval is short relative to the harm. A failure detectable only after months of
accumulated cases is not controlled by monitoring where the harm from each instance is
immediate and irreversible.
3. An automatic mitigation exists. On signal exceedance the device reverts to a defined safe state,
such as human interpretation, without awaiting manual action by the manufacturer.
4. Output is reproducible. Where output varies across repeated runs on identical inputs, monitored
performance cannot reliably be attributed to the device configuration under review, and rebenchmarking against a premarket baseline loses its footing.
For functions at the higher end of the risk profile described in Section III — in particular those that
generate diagnoses for serious or critical conditions without evaluation by a qualified person before
effect — these conditions will rarely hold together. A presentation too uncommon to appear in the
monitored population is not characterized by monitoring it, and the harm from a missed time-critical
finding occurs before any signal could be observed. In answer to the final part of Question 18, these are
the device types for which we believe the approach would not be appropriate.
Postmarket monitoring should instead build on premarket evidence designed to be as complete as the
device permits, and that is in large part a matter of design. Sponsors can bound the input and output
space, control the variability of output, and design premarket evaluation to capture as much of the
admitted space as practicable. We do not regard incomplete premarket characterization as an inherent
property of GenAI-enabled devices, and we would not want the framework to treat it as one.
B. Why premarket evidence cannot fully characterize deployed performance
Even premarket evidence designed as well as possible will not characterize every aspect of a device’s
performance in deployment. We identify four reasons. Each is a proper object of postmarket monitoring,
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 22 of 45
Docket No. FDA-2026-N-7874 · Contents
and each is best addressed by one of the approaches the discussion paper describes, which we
support:
Reason Why premarket evidence is incomplete Monitoring approach best
suited
Model evolution The device or an underlying model Periodic device
changes after authorization, through benchmarking, triggered by
sponsor modification, continuous defined changes
adaptation, or third-party foundation
model updates
Output variability Output that varies across runs may vary Periodic sample-based
in ways a finite premarket sample did clinician review
not reveal
Input and population Equipment, protocols, practice patterns, Performance degradation
drift and populations change after monitoring
authorization
Open-ended input Premarket evaluation samples only part Periodic sample-based
and output of the space an open-ended device clinician review, directed at
admits unanticipated inputs and
outputs
The first reason encompasses each of the three kinds of change the discussion paper describes:
intentional modifications, model-evolution changes, and unplanned changes arising from third-party
foundation models. The second and fourth are the two properties comprising the fourth dimension of
risk proposed in Section III. The same design characteristics that raise the premarket evidentiary burden
therefore also drive the need for postmarket monitoring, which is a further reason to address them by
design.
C. Scaling monitoring to the reasons that apply
The discussion paper anticipates that the scope, frequency, and level of evidence of postmarket
monitoring “might vary in proportion to the risk profile of the device.” We agree, and recommend that
the intensity of monitoring be scaled specifically to which of the four reasons above apply to a given
device.
On output variability, we wish to record a technical point we believe is important for the Agency. Setting
decoding temperature to zero does not by itself guarantee reproducible output in a production serving
environment. Because batch size and composition vary with concurrent load, and because common
inference kernels are not invariant to batch composition, the order of floating-point reductions changes
with the batch, and the output can change with it. Contributing factors include floating-point nonassociativity in reduction operations, dynamic batching, expert routing that depends on batch
composition, speculative decoding, and hardware and driver variation. 35 We therefore recommend that
any reproducibility expectation be framed behaviorally — as demonstrated identical output across
repeated runs on identical inputs under the deployed configuration — rather than as a configuration
parameter that can be asserted but not verified.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 23 of 45
Docket No. FDA-2026-N-7874 · Contents
We want to be precise about the contrast with conventional artificial intelligence, and to avoid
overstating it. Conventional inference with fixed weights is deterministic when explicitly configured for
determinism: pinned library and driver versions, deterministic algorithm selection, and deterministic
operations. It is not automatically deterministic, and non-deterministic accumulation and automatic
algorithm selection are documented sources of variation. The accurate contrast is that these are
configurable properties for conventional inference, whereas the batch-composition dependence
described above persists in generative serving irrespective of determinism settings.
A device with fixed weights, reproducible output, and a bounded input and output space is exposed
principally to input and population drift. A device with variable output, built on a third-party foundation
model and admitting open-ended input, is exposed to all four. We recommend that monitoring cadence,
scope, and re-benchmarking triggers differ substantially between these cases.
We want to be equally clear about what reproducibility does not do. It does not remove the need for
postmarket monitoring. Equipment replacement, protocol change, contrast-agent change, and shifts in
the referred population all degrade the real-world performance of a device whose weights have not
changed at all, and no degree of reproducibility will detect that. We are not proposing that any category
of device be exempted from monitoring. We are proposing that the intensity of monitoring be
proportionate to the sources of variation present, which we believe is both safer and more consistent
with least burdensome principles than a uniform expectation.
Where premarket evidence characterizes performance of a conventional AI device across a broad and
representative range of acquisition conditions, equipment, and populations, we believe the scope and
cadence of proactive performance monitoring may reasonably be reduced.
D. Machine-based supervisory monitoring
With respect to Question 20, we support the concept of machine-based supervisory monitoring, and
note that it is already in use. An AI device for autonomous triage in organized breast cancer screening,
certified in the European Union in September 2026, relies on a supervisory layer that sets and monitors
each site’s operating point, tracks changes in equipment, system health, and daily performance signals,
and returns a site to full radiologist reading when those signals move outside defined limits. 36
Question 20 also asks about the evaluation and reliability of the supervisory agent itself. We
recommend that a supervisory function of this kind be treated as a device software function in its own
right, with its own requirements specification, verification, validation, and performance characterization,
rather than as telemetry or a quality-system artifact. The reason is that it is a risk control measure: it
changes the device’s behavior, for example by withdrawing autonomous operation and returning cases
to human interpretation, and its reliability therefore bears directly on patient safety. Consistent with
Section V.A, we regard such a function as complementing premarket evidence, not as a basis for
reducing it.
E. Devices built on third-party foundation models
With respect to Question 24, we agree that devices built on third-party foundation models present a
change-control problem of a different character, because the change originates outside the
manufacturer’s quality system. Rather than a categorical restriction, we recommend a combination of
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 24 of 45
Docket No. FDA-2026-N-7874 · Contents
controls: version pinning as a device specification rather than a deployment convenience; contractual
change-notification commitments from the model provider; mandatory re-benchmarking against the
premarket baseline upon any change to the underlying model; and use of a Predetermined Change
Control Plan as the mechanism through which anticipated changes and their acceptance criteria are
agreed in advance.37 We would regard a device whose clinical output depends on a model that may
change without the manufacturer’s knowledge, and without re-benchmarking, as not adequately
controlled.
With respect to Question 25, we agree with the discussion paper that submission of a Foundation Model
Master File “would not constitute authorization of the underlying model for any device intended use,”
and that sponsors would remain responsible for demonstrating the safety and effectiveness of their
own device. Such files could usefully inform review. They should not substitute for device-level
evidence.
VI. Architectural Safeguards for High-Risk-Pro2le Device
Functions, Including Autonomous Functions
RESPONSIVE TO QUESTIONS 4, 11, 18 AND 26.
A. Functions whose output is not reviewed before it takes e@ect
The discussion paper anticipates functions whose output is not reviewed by a qualified clinician before
it takes effect. It places fully autonomous operation at the far end of its activity axis; it addresses
patient-facing and generalist-facing functions; and its clinical confirmation approaches — retrospective
evaluation on real patient inputs, shadow deployment, standardized patient interactions, and clinician
adjudication of real cases — are well suited to evaluating standalone performance in those settings. The
paper also recognizes that “GenAI-enabled devices interact with users, workflows, and patient
populations in ways that may not be captured by structured benchmarking.” It does not, however,
address how the risk of such functions might be controlled by design. We offer the following as a
constructive contribution.
B. Risk as a continuum
We do not regard autonomous devices as a categorically distinct class. Risk is continuous across the
four dimensions described in Section III, and a supervised function can carry risk comparable to an
autonomous one. A function that is stochastic, admits open-ended input and output, generates
diagnoses, and has a basis its supervising clinician cannot evaluate, places a supervisor in the position
of verifying claims whose origin they cannot inspect. The evidence on automation bias shows that such
verification is imperfect,15 and the discussion paper itself identifies the risk that a user accepts an output
“without appropriate scrutiny because of the device’s fluency or perceived authority.”
The central regulatory difficulty with such functions is not that they may be inaccurate; accuracy is
measurable. The difficulty is that the set of inputs on which they may fail cannot be exhaustively
enumerated, and a control cannot be specified against a failure mode that cannot be described. That
difficulty applies to autonomous functions, and it applies equally to generative functions with
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 25 of 45
Docket No. FDA-2026-N-7874 · Contents
unbounded input and output operating under supervision. We therefore recommend that the safeguards
below attach to the risk profile of a function, and not to whether it is labeled autonomous.
C. Three safeguards
We suggest that three architectural properties, taken together, convert an unbounded failure space into
a bounded and measurable one. They are properties of a class of designs rather than of any particular
implementation, and could be expressed as controls applicable to high-risk-profile functions generally.
1. Quantified uncertainty at each processing stage. Each stage of processing emits a quantified
measure of uncertainty appropriate to its task, and each measure is validated by demonstrating
that it is associated with error, rather than assumed to be. Where downstream processing is
reproducible, uncertainty at an earlier stage can be propagated forward to show whether the final
clinical conclusion is stable under the uncertainty present in the inputs that produced it. Selfreported model confidence has documented limits of calibration under distribution shift, 38 which
is why we emphasize validation against observed error.
Example. Consider a function that assesses the quality of a wearable electrocardiogram signal, detects
individual beats, and classifies the cardiac rhythm. A segment degraded by motion yields high
uncertainty at the signal-quality stage. When that uncertainty is carried forward, the rhythm
classification may be shown to change under small, clinically irrelevant perturbations of the degraded
segment. The function can then decline to report the classification, rather than reporting a
determination whose basis is unreliable.
2. An independent secondary evaluation path (safety-net).39 A second path evaluates each case,
structurally independent of the primary path in its task formulation, its training data, and its
architecture. Independence should be established by auditable design documentation rather than
asserted; agreement between genuinely independent paths is then a performance result, not a
design assumption. This safeguard addresses a failure mode that quantified uncertainty alone
cannot detect: a single model producing an incorrect output with high confidence.
Example. Consider a function that classifies skin lesions from photographs. A secondary path, trained
for a different task — determining only whether any lesion warranting clinical evaluation is present — on
different data and with a different architecture, evaluates the same image. Where the primary path
characterizes a lesion as benign with high confidence, but the secondary path indicates that a lesion
warranting evaluation is present, the case is not finalized.
3. Structured deferral. The function declines to finalize, and routes the case for human
interpretation, when any of the following occurs: a validity precondition fails; a finding is
identified that cannot be characterized within the authorized output set; the independent paths
disagree materially; or system-level uncertainty exceeds a prespecified threshold calibrated to a
stated maximum error rate.39
Examples. An image acquired on equipment outside the validated specification, or of inadequate
quality, is deferred before any determination is reported. An abnormality that is detected but falls
outside the conditions the function is authorized to characterize is deferred, rather than omitted from an
otherwise complete-appearing output. We note that the first FDA-authorized autonomous diagnostic
device incorporates a determination of examination quality alongside its diagnostic determination. 31
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 26 of 45
Docket No. FDA-2026-N-7874 · Contents
D. Safeguards proportionate to risk pro2le
These safeguards are not an all-or-nothing requirement. We recommend that the extent to which they
are expected scale with the risk profile of the function:
Risk profile Illustrative characteristics Safeguards expected
Lower Presenting functions; or generating functions Conventional validation;
for non-serious conditions with bounded, faithfulness benchmarking
reproducible output, received by a user able to where presenting
evaluate it
Moderate Generating measurements or detections for Quantified uncertainty at
serious conditions, with bounded output, each stage, with structured
received by a clinician able to evaluate it deferral
Higher Generating diagnoses or management for All three safeguards,
serious or critical conditions; or variable or including an independent
open-ended output; or no qualified evaluator secondary evaluation path
before effect, including autonomous, patientfacing, and generalist-facing functions for
specialist content
E. Performance characterization
The performance characterization that follows from this architecture is tractable and specifiable in a
way that whole-device accuracy alone is not. It comprises the error rate among cases the function
finalizes, reported with confidence intervals and stratified by clinical consequence — since a tolerance
appropriate for an incidental finding is not appropriate for a time-critical one — together with the
deferral rate, reported separately as an operating parameter rather than as a safety endpoint.
We note and support the paper’s treatment of both under-deferral and over-deferral as failures under
element S.3. A function that defers excessively provides little benefit and may erode clinical trust; a
function that defers too rarely is not meaningfully safeguarded. Both should be reported, and the
operating point at which a sponsor sets the trade-off should be prespecified and justified.
F. Two illustrations
A supervised function with a high risk profile. Consider a generative function that receives an
abdominal CT examination with a free-text prompt, and returns an unrestricted narrative report that may
include diagnoses, management recommendations, and image annotations, across an unbounded set of
possible conditions, using sampled generation. Every output is reviewed by a radiologist. The function
nonetheless generates diagnostic and management claims; it addresses conditions up to the critical; its
output is variable and its input and output space unbounded; and the basis for its claims cannot be
evaluated by the radiologist who reviews it. Supervision does not remove these characteristics. In our
view this function warrants the full set of safeguards described above, notwithstanding that it is
supervised.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 27 of 45
Docket No. FDA-2026-N-7874 · Contents
A paired contrast. Consider two functions that each detect and characterize fractures on wrist
radiographs and deliver the result directly to an emergency physician, without radiologist review. The
significance of the output, the state of the condition, and the user are identical. The first produces its
output from a stochastic model with open-ended output, whose basis the physician cannot evaluate.
The second is composed of separately verifiable processing stages, produces reproducible output
within a bounded set of claims, and makes its intermediate results available to the physician. Holding
the other three dimensions constant, the technological characteristics alone produce a substantially
different risk profile, and a correspondingly different expectation of evidence. The second design is
intended to support evaluation of the basis for its output; whether it does so for its intended user is
itself a claim to be demonstrated, as discussed in Section III.C.3.
VII. Applicability Beyond Generative Architectures
RESPONSIVE TO QUESTIONS 1, 17 AND 26.
A. Several of these recommendations are technology-agnostic
The distinction between generating and presenting clinical-decision content, the dimensions of risk
proposed in Section III, the expectation that the input and output space be bounded, claim-level
evaluation, the principle that the comparator follow the standard of care of the intended environment,
the architectural safeguards described in Section VI, and the scaling of monitoring to the reasons
premarket evidence is incomplete are all independent of whether a device is generative. We
recommend that CDRH state this explicitly, so that a framework developed in response to generative
architectures does not become inapplicable to functions that raise the same questions by other means.
Several questions the discussion paper raises for GenAI-enabled devices — in particular the role of the
user, and the evaluation of functions whose output is not reviewed by a qualified clinician before it
takes effect — are equally open for conventional AI-enabled devices, and are not at present addressed
by FDA guidance for them.
B. Autonomy is not a generative-AI question
The risk of autonomous operation arises from the absence of qualified human review before an output
takes effect, not from the class of model that produced the output. The first FDA authorization of an
autonomous diagnostic device involved a non-generative architecture.31 If CDRH develops its thinking
on autonomous operation only within the generative track, device functions that are autonomous but
not generative will be either unaddressed or accommodated within a framework built around a different
failure mode. We recommend that autonomy be treated as a cross-cutting topic in its own right.
C. A single risk framework
With the fourth dimension proposed in Section III, a single framework can serve generative and
conventional device software functions alike. Conventional functions with fixed weights, reproducible
output, and bounded inputs and outputs will generally sit at the lower end of that dimension; generative
functions with open-ended input and output and variable output will sit at the higher end. For the same
intended use, the latter will warrant substantially more evidence, and the framework will say so for a
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 28 of 45
Docket No. FDA-2026-N-7874 · Contents
stated reason rather than because of the technology’s label. We believe this is the approach most
consistent with the discussion paper’s own statement that FDA “does not regulate GenAI as such”: the
framework turns on the characteristics that bear on safety and effectiveness, and applies equally to any
device that has them.
D. Why generative implementations warrant additional scrutiny, and the limits of
that argument
Relative to a conventional implementation performing the same function, a generative implementation
differs in three respects that bear on evaluation. Output reproducibility must be engineered and
demonstrated rather than following from fixed weights and a deterministic configuration. The output
space is bounded only by deliberate design rather than by the structure of the model. And established
evaluation methodology, existing special controls, and accumulated review experience apply directly to
the conventional case and only partially to the generative one.
The emerging comparative evidence is consistent with this. A recent diagnostic study comparing
general-purpose vision-language models against task-specific models on an identical clinical task
reported substantially lower performance for the general-purpose systems, 40 and a systems-level
review argues that such models are most plausibly valuable as a semantic layer consuming structured
outputs from task-specific models, rather than as replacements for detection, segmentation, and
quantitative measurement.41 Reported reproducibility of generative models is finding-dependent, with
salient findings reproducing reliably across runs and more nuanced findings substantially less so. 42
We want to state the limit of this argument plainly. None of it establishes that conventional artificial
intelligence is safe. A deterministic model can be confidently and consistently wrong, and a
reproducible error remains an error. The advantage of conventional implementations is evaluability, not
inherent safety. Our recommendation is accordingly not that generative devices be disfavored, but that
the evidence required be matched to the properties of the device.
E. Classi2cation continuity and substantial equivalence
Intended use governs classification. A generative function that produces a quantitative measurement, a
detection, or a diagnosis has the same intended use as a conventional device performing that function,
and in our view should remain within the existing device type for that intended use. We do not
recommend that CDRH create parallel classifications for generative implementations.
The question that requires attention is substantial equivalence. Section 513(i)(1)(A) of the FD&C Act
provides that a device with the same intended use as its predicate but different technological
characteristics is substantially equivalent only where it is shown to be as safe and effective as the
predicate and does not raise different questions of safety and effectiveness. Under FDA’s established
approach, technological characteristics that raise different questions of safety and effectiveness lead to
a determination that the device is not substantially equivalent.43
We do not recommend that generative implementations be categorically excluded from substantial
equivalence to conventional predicates. That would regulate GenAI as such, and it would disregard the
fact that output variability and open-endedness can be controlled by design. We recommend instead
that the determination turn on the characteristics themselves. A generative implementation whose
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 29 of 45
Docket No. FDA-2026-N-7874 · Contents
output is not reproducible, or whose input and output space is open-ended, relative to a predicate
whose output is reproducible and whose space is bounded, raises a different question of safety and
effectiveness — one the predicate’s data cannot answer — and in our view should not be found
substantially equivalent to that predicate. A generative implementation that has demonstrably controlled
both may not raise such a question.
Three examples illustrate the point.
• A generative function returning a volumetric measurement. Its intended use matches that of a
cleared quantitative imaging device. If repeated analysis of identical images yields different
values, the predicate’s repeatability data does not characterize this device, and the variability of
its output is a different question of safety and effectiveness that the predicate cannot answer.
• A generative function producing lesion detections. Its intended use matches that of a cleared
computer-assisted detection device. The evaluation of that device type presumes a fixed
operating point, and a detection set that varies across runs is not characterized by a single
receiver operating characteristic curve.
• A generative narrative added to an otherwise conventional analysis. Under the multiple
function framework, the conventional function may well be substantially equivalent. The
narrative-generation function is a new function requiring its own evaluation, and the effect of the
new function on the performance of the cleared function must be assessed.
This does not require new classification regulations. It requires that the existing substantial equivalence
standard be applied to technological characteristics that are new.
VIII. Generalizability: Synthetic Data and Independent
Benchmarking
RESPONSIVE TO QUESTIONS 10, 12, 13 AND 16.
The discussion paper identifies generalizability as a benchmarking element in its own right: whether a
device’s safety and clinical proficiency properties “hold consistently across foreseeable variation in
inputs, runtime conditions, and conversational contexts” (R.1), and “across clinically relevant
populations, including populations that may be underrepresented in training data” (R.2). 1 In its
discussion of clinical confirmation, CDRH considers that synthetic data “could potentially supplement
evaluation” where subgroups are small or hard to sample; and in Section V.D.2 of the paper it asks
whether qualified, independent third parties might contribute to benchmarking, for example “by
providing larger or sequestered evaluation datasets” or by “offering independence between
development and testing.” We strongly support all three directions, and we believe they belong
together: generalizability is where the evidence for AI-enabled devices most needs strengthening,
validated synthetic data is the most direct way to strengthen it, and independent third parties are well
placed to provide such data. As with most of our comments, these considerations apply equally to
conventional AI-enabled devices.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 30 of 45
Docket No. FDA-2026-N-7874 · Contents
A. Why generalizability deserves greater weight
As AI-enabled devices become more accurate on average, the clinically meaningful question shifts from
how often a device is right to where its remaining errors are concentrated. Aggregate performance can
conceal clinically important failures confined to a subset of cases,44 and models that perform
indistinguishably on a held-out test set can behave very differently when their inputs shift in realistic
ways.45 As benchmarks approach saturation (Question 10), the evidence that most distinguishes one
device from another increasingly lies in how consistently it performs across the conditions of its
intended use.
A real-world test set, however carefully assembled, is a sample of the intended-use space: the
combinations of condition, presentation, acquisition equipment and protocol, and patient population in
which a device will be used. A test set of realistic size populates only part of that space. Five hundred
cases divided evenly across eighteen combinations — for example, three equipment vendors, two
acquisition protocols, and three age groups — leave fewer than thirty cases in each. A device that
detects 13 of 14 positive cases in one such combination has an observed sensitivity of 92.9 percent,
with a 95 percent confidence interval extending from below 70 percent to above 98 percent. The
limitation is not confined to rare diseases: common presentations on less common equipment, or in less
common populations, are sampled just as thinly. Published analyses of authorized devices reflect the
same constraint. Of 903 FDA-authorized AI-enabled devices in one analysis, clinical performance
studies were reported for 505, and fewer than a third of those provided sex-specific data. 46
Real-world data and data collected outside the United States can enlarge a test set, and we support
their use as the paper proposes. They do not, however, change its structure. Additional real-world
cases arrive in proportion to how often each combination occurs in practice, so thinly sampled
combinations remain thin; and data from other health systems bring differences in population, practice,
and equipment whose relevance to the intended use must be established. Real-world data also vary
many factors at once — site, equipment, protocol, and population tend to change together — so that
when performance falls, the data cannot always show why. In one widely cited study, models trained to
detect pneumonia on chest radiographs could identify the hospital system from which a radiograph
came with near-perfect accuracy, and performed better on internal than on external data in three of five
comparisons.47
B. What synthetic data contributes
Synthetic data addresses these limitations directly. Its value for evaluating generalizability rests on six
properties:
• Coverage by design. Cases can be generated to populate every combination the intended use
specifies, in numbers set by the precision required rather than by how often each combination
occurs in practice — including the small and hard-to-sample subgroups the paper identifies.
• Controlled variation. One factor — equipment, protocol, image quality, or patient characteristics
— can be varied while the anatomy and the finding are held constant, so that the effect of each
factor on performance is measured directly rather than inferred.
• Known reference standard. Each case is generated from a specified configuration, so its
reference standard is known by construction, free of inter-reader variability, and scoring can be
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 31 of 45
Docket No. FDA-2026-N-7874 · Contents
fully automated and reproducible. Combined with the claim-level evaluation described in Section
IV.D above, this extends to generative functions: a generated report can be scored claim by claim
against findings that are known exactly.
• Freedom from contamination. A new test set can be generated for each evaluation. Its cases
cannot have appeared in any device’s training data and cannot be learned through repeated use,
which addresses the contamination and saturation concerns raised in Question 10.
• Speed and repeatability. Evaluation can be automated end to end and repeated whenever
needed — before a submission, at each modification under a predetermined change control
plan,37 and periodically after marketing — without the months required to assemble and
adjudicate a new clinical dataset. The paper anticipates this, noting that synthetic data generation
and simulation tools permit testing “in a manner that is predictable, standardized, and amenable
to automation.”1
• Testing at the boundary. Inputs can be constructed just outside the conditions for which a
device is validated. This is the most direct way to verify that a device recognizes such inputs and
defers rather than reports — the behavior that Sections IV.B and VI above identify as essential.
In answer to the first part of Question 13, we believe synthetic data is particularly well suited where the
process that produces a device’s input is well understood and can be modeled — as in medical imaging,
where image formation follows known physics — and to functions whose output can be scored against
a known reference: detection, localization, measurement, and classification. FDA’s own VICTRE project
demonstrated the approach: an in silico imaging trial of approximately 3,000 computational breast
models agreed with a clinical trial in finding that digital breast tomosynthesis improves lesion
detectability over digital mammography.48 Synthetic data is less suited where the reference standard
depends on information a generator does not represent, such as clinical outcomes, and to the study of
how clinicians use a device, which belongs to clinical confirmation.
C. Establishing that a synthetic benchmark is 2t for purpose
A synthetic benchmark earns reliance through evidence of its own validity, and FDA has already
published a framework on which such evidence can be built: the risk-informed credibility assessment
described in its guidance on computational modeling and simulation.4 We suggest that this evidence
show three things: that synthetic cases are realistic across the intended-use space, including by
blinded expert discrimination from real cases; that each labeled finding is perceptible in the generated
input, so that the known reference standard is also a fair one; and that performance measured on the
benchmark agrees with performance measured on real-world data, for reference devices, wherever
both are available. The third provides the evidence Question 10 asks for — that benchmark performance
predicts real-world behavior — and once it has been shown where both kinds of data exist, the
benchmark can be relied on where real-world data are thin. For the same reason, and in answer to
Question 12, we suggest that synthetic and real-world results be reported separately by default, and
combined into a single estimate only under a pre-specified method once that agreement has been
shown.
The concern raised in the second part of Question 13 — that synthetic data generated by models of the
same class as the device could reproduce the very gaps an evaluation is meant to detect — is
addressed most directly by independence. A generator developed independently of the device under
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 32 of 45
Docket No. FDA-2026-N-7874 · Contents
evaluation, for example one built on explicit models of anatomy, pathology, and image formation, as
VICTRE was, is unlikely to share the device’s blind spots. An independent third-party benchmark
provides that independence by design.
Synthetic data complements real-world evidence; it does not replace it. Real-world data remain the
direct measure of performance in the deployed population and capture variation that no generator
anticipates, while synthetic data extends evaluation to what real-world data cannot practically reach.
Where a condition cannot be evidenced at the required level by either, the approach described in
Section IV.C above — exclusion from the authorized output set and deferral to human interpretation —
continues to apply.
D. The role of independent third parties
We believe independent third parties can contribute substantially to benchmarking, and that synthetic
benchmarking is where their contribution would be greatest, for three reasons:
• Capability. A validated synthetic benchmark requires generators, reference data, and validation
evidence that many sponsors, particularly smaller ones, cannot build or assemble in-house.
Third-party provision makes rigorous evaluation of generalizability available to every sponsor,
rather than only to those able to build that capability themselves.
• Independence. A benchmark developed and operated independently of the device under
evaluation offers the “independence between development and testing” the paper contemplates,
and the safeguard Question 13 seeks.
• Comparability. Devices evaluated on the same benchmark under the same protocol can be
compared directly: successive versions of one device, a device and its predicate, and devices
with the same intended use. This is difficult today, when each sponsor evaluates its device on a
dataset of its own.
On the independence criteria that Question 16 raises, we suggest that independence be defined in
relation to the device under evaluation, consistent with the paper’s formulation for expert adjudicators,
who would be “structurally independent from the device sponsor.”1 Those best placed to build
evaluation tools are often those who develop devices, because they encounter the gaps in available
evidence first. FDA’s MDDT program accepts qualification proposals from device sponsors as well as
from other tool developers,49 and FDA recently contracted with a developer of AI radiology devices to
test a new approach to evaluating generative AI in radiology.50
We further suggest that confidence in a third-party evaluation rest on how it is conducted, not only on
who conducts it. An evaluation with the following features is verifiable by every party:
• Pre-specified. The protocol, metrics, and acceptance criteria are fixed and provided to FDA
before the device is evaluated, together with a verifiable, time-stamped record of the exact test
set, such as a cryptographic hash, so that neither can be altered once results are known.
• Automatically scored. Results are scored against the known reference standard, leaving no room
for discretionary judgment in either direction.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 33 of 45
Docket No. FDA-2026-N-7874 · Contents
• Identically reported. FDA and the sponsor receive the same report at the same time, with results
broken down across the intended-use space and illustrative cases in which the device failed, so
that both can see where and how it fails.
• Fully auditable. FDA has confidential access to the complete test set, the scoring method, and
documentation of the generator — for example through the existing Device Master File program,
which the paper proposes to leverage for foundation models (Section VII.A). The test set itself is
not released to sponsors, and cases disclosed in reporting are retired from future use.
This model is established outside medicine. NIST’s face recognition technology evaluation runs
continuously, with developers submitting algorithms for testing rather than receiving the test data, and
with results reported for each algorithm.51
The role of a third party is to measure; acceptance criteria and the determination of safety and
effectiveness remain with FDA and the sponsor. To preserve competition and innovation, we suggest
that FDA qualify tools rather than designate providers, that access be offered on published, nondiscriminatory terms, and that sponsors remain free to provide equivalent evidence by other means.
E. Recommendations
We respectfully recommend that FDA:
1. Retain generalizability as a distinct element of the evidence expected for AI-enabled devices,
generative and conventional alike, with performance reported across the relevant strata of the
intended use rather than in aggregate alone.
2. Recognize validated synthetic benchmarks as an acceptable source of evidence of
generalizability, including subgroup performance and robustness to acquisition conditions, in
premarket submissions, in the verification of modifications under a predetermined change control
plan, and in postmarket re-benchmarking.
3. Publish expectations for qualifying synthetic benchmarking tools through the MDDT
program,52 including the evidence of realism, fairness of the reference standard, and agreement
with real-world performance described above, and give priority to proposals that address the
evidence gaps the discussion paper identifies.
4. Continue to support the development of independent benchmarking capability through its
existing research and qualification mechanisms.
5. Reflect such evidence in postmarket expectations: where premarket evidence characterizes the
performance of a conventional AI device across a broad and representative range of acquisition
conditions, equipment, and populations, the scope and cadence of proactive performance
monitoring may reasonably be reduced, as proposed in Section V.C above.
IX. Conclusion
We appreciate CDRH’s decision to engage stakeholders at this stage rather than after a framework has
been settled. We think it is the right decision, and we have tried to respond in kind. Several of the
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 34 of 45
Docket No. FDA-2026-N-7874 · Contents
recommendations above would raise rather than lower the evidentiary expectations for the class of
devices we work on.
We hold these views because we believe the principal risk to this field is not that regulation will be too
demanding. It is that an early and avoidable failure in a high-consequence clinical setting would set
back the deployment of technologies that are, in many parts of the world and increasingly in parts of the
United States, the only realistic path to timely specialist-level interpretation. A framework that is precise
about what must be demonstrated, and proportionate about what need not be, serves patients and
sponsors alike.
We would welcome the opportunity to discuss any of the above with CDRH staff, and would be glad to
contribute to standards development, benchmark design, or workshop discussion on the evaluation of
autonomous and imaging-based device functions.
Respectfully submitted,
Reza Rahmanzadeh, MD, PhD
CEO & Founder @ UltraAI
Reza@theultra.ai
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 35 of 45
Docket No. FDA-2026-N-7874 · Contents
Appendix A. Generating and Presenting: Illustrative
Examples
The following examples illustrate how the distinction between presenting and generating clinicaldecision content may be applied. They are illustrative rather than exhaustive. In each case the
characterization applies to individual claims within an output, not to the output as a whole.
Example A1. Drafting a radiology impression
Input. The findings section of a report written by a radiologist.
Presenting. The function drafts an impression that restates, in condensed form, the findings the
radiologist documented. Each clinical statement in the impression is traceable to a statement in the
findings.
Generating. The function adds a statement that is not traceable to the radiologist’s findings or to an
identified external source — for example, that a finding is suspicious for a diagnosis the radiologist did
not state, or that follow-up imaging is advisable at an interval the radiologist did not specify. That
statement is a clinical claim originating with the function.
Implication. The generated statement warrants validation of the function’s performance in making that
category of claim. The restated findings warrant faithfulness benchmarking.
Example A2. Summarizing serial oncologic imaging reports
Input. Prior radiology reports documenting the measurements of a brain tumor over time.
Presenting. The function summarizes the prior reports, including the measurements each documents.
Generating. The function adds that the treatment has not been effective, and that the patient should
receive further radiotherapy. The first is a diagnostic claim and the second a management claim; both
originate with the function.
Presenting, with published criteria applied. The function states the documented change in
measurements; states that this change meets the published definition of progression under named
response assessment criteria;53 and states what a named treatment guideline specifies for patients
meeting that definition. The characterization of progression and the associated management statement
are drawn from identified sources and applied to measurements documented by a radiologist.
Note. If the function derived the tumor measurements from the images rather than from the reports,
those measurements would be generated quantitative measurements, even where the characterization
that follows is presented.
Example A3. A laboratory value and a published guideline
Input. A patient’s LDL cholesterol value.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 36 of 45
Docket No. FDA-2026-N-7874 · Contents
Presenting. The function compares the value against the threshold given in a named lipid-management
guideline, and conveys the classification and associated statement that the guideline provides for that
range, with a reference to the guideline and its version. The claim is drawn from the identified source.
Generating, where the output may appear identical. The function submits the same value to a
general-purpose large language model, which returns text of the form “this level may increase the risk
of myocardial infarction by X percent over Y years, according to [guideline] [reference].” Even where
the model has been instructed to cite the guideline, and even where the output is textually identical to
the presenting case, the claim was produced from the model’s learned parameters. It cannot be
established from the output that the cited guideline contains the figure given, that the version
referenced is current, or that the reference corresponds to the claim.
Implication. This is the case in which the distinction matters most. The presence of a citation does not
establish that a claim is presented. A sponsor characterizing a function as presenting should be able to
demonstrate that its claims are drawn from the identified source.
Example A4. Synthesizing reports, laboratory results, and clinical documentation
Input. A radiology report documenting longitudinally extensive transverse myelitis; a laboratory result
documenting aquaporin-4 IgG seropositivity; and clinical documentation of optic neuritis.
Output. “The documented findings fulfil the international consensus diagnostic criteria for
neuromyelitis optica spectrum disorder.”54
Presenting. The function applies the published criteria to documented findings, each of which
originates with a qualified clinician or laboratory. The diagnostic characterization is drawn from the
published criteria; the function has not originated a diagnostic claim of its own.
Generating. The same inputs are submitted to a generative model, which produces the same sentence.
The diagnostic claim now originates with the model. Whether the model’s parameters reflect the criteria,
which version of them, and how reliably, cannot be established from the output.
Note on faithfulness. The published criteria require, among other conditions, the exclusion of
alternative diagnoses. A presenting function that states the criteria are fulfilled without conveying that
condition has misrepresented its source, even though it has not generated a diagnosis of its own. This
is why presenting functions require faithfulness benchmarking.
Example A5. A wearable device combining physiological and health data
Input. Continuous physiological signals from a wearable device, together with other health information
about the wearer.
Generating. The function combines these inputs to produce a diagnostic classification, an alert, or a
risk score. Each is a clinical claim originating with the function, and each warrants validation of the
function’s performance in producing it.
Note. A physiological measurement computed from a raw sensor signal — for example, heart rate or
oxygen saturation derived from an optical signal — is itself a generated quantitative measurement. A
single output from such a device may therefore contain claims at several levels of the significance axis.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 37 of 45
Docket No. FDA-2026-N-7874 · Contents
Example A6. Renal dose adjustment
Input. A patient’s estimated glomerular filtration rate and a prescribed medication.
Presenting. The function returns the dose adjustment that the medication’s prescribing information
specifies for that range of renal function, identifying the prescribing information and its version.
Generating. The function determines a dose from the patient’s overall clinical picture, drawing on its
learned parameters.
Implication. This is the kind of quantitative task the discussion paper describes under element E.3.
Where the function presents, faithfulness to the prescribing information is the appropriate evidence.
Where it generates, the dose it produces is a clinical-management claim, requiring validation of the
function’s performance in making it.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 38 of 45
Docket No. FDA-2026-N-7874 · Contents
Appendix B. Risk Pro2les: Illustrative Examples
The following examples illustrate how the four dimensions proposed in Section III combine into a risk
profile, and how the profile would inform the evidence and safeguards expected. They are illustrative
and are not intended to characterize any particular product.
Example B1. Summarizing a medication history for a clinical pharmacist
Function. Summarizes the medications recorded across a patient’s records, each entry linked to the
record from which it was drawn.
Significance. Presenting.
Condition. The function does not itself address a condition, although an omitted medication could
contribute to harm.
User. A clinical pharmacist, who can verify each entry against its linked source.
Characterizability. Output constrained to extracted entries, each linked to a source.
Profile and evidence. Lower. Faithfulness benchmarking, including specifically the omission of material
entries.
Example B2. Notifying a wearer of possible atrial 2brillation
Function. Analyzes a single-lead electrocardiogram from a wearable device and notifies the wearer of
possible atrial fibrillation.
Significance. Generating a detection.
Condition. Serious.
User. The wearer, who cannot evaluate the output or its basis; applicability is clear where the
instruction is to seek clinical evaluation.
Characterizability. Reproducible output within a defined set of rhythm classes, from a defined input.
Profile and evidence. Moderate. Quantified uncertainty with structured deferral — for example,
reporting an inconclusive result for a degraded signal rather than a classification.
Example B3. Early warning of clinical deterioration in intensive care
Function. Computes a score estimating the risk of clinical deterioration from vital signs and laboratory
values, and alerts the care team.
Significance. Generating a detection.
Condition. Critical.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 39 of 45
Docket No. FDA-2026-N-7874 · Contents
User. An intensivist, who can partly evaluate the output against the clinical picture, but typically not its
basis; applicability is good, since the alert enters an established pathway.
Characterizability. Reproducible output from defined inputs.
Profile and evidence. Moderate to higher. Validation of alert performance at the operating thresholds
actually used; quantified uncertainty communicated with each score. This is the kind of risk score the
discussion paper treats as non-directive; at high values it functions as a direction to act, which is why
we propose grading by significance rather than by directiveness.
Example B4. A consumer application assessing skin lesions from photographs
Function. Assesses a photograph of a skin lesion taken by the user and indicates whether clinical
evaluation is advisable.
Significance. Generating a diagnosis.
Condition. Potentially critical, since a missed melanoma may be.
User. The user, who cannot evaluate the output or its basis; applicability depends on the clarity of the
recommended next step.
Characterizability. If implemented with a stochastic model producing open-ended output: high. If
implemented with a closed output set and an image-quality gate: lower.
Profile and evidence. Higher in either implementation, because the condition and the user keep the
profile high; all three safeguards. Bounding the output reduces the characterizability burden but does
not change the other dimensions.
Example B5. A conversational assistant advising patients whether to seek
emergency care
Function. Converses with a patient about symptoms such as chest pain and advises whether to seek
emergency care.
Significance. Generating a clinical-management determination.
Condition. Up to critical.
User. The patient.
Characterizability. Open-ended conversational input and variable output: high. A bounded
implementation — structured symptom intake, a closed set of dispositions, and mandatory escalation
for defined features — would be lower on this dimension.
Profile and evidence. Higher. All three safeguards; under-escalation and over-escalation evaluated as
sensitivity and specificity for each condition, as proposed in Section III.E.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 40 of 45
Docket No. FDA-2026-N-7874 · Contents
Appendix C. Bounding and Claim-Level Evaluation:
Illustrative Examples
C1. Bounding the input and output space by intended use
Intended use Input bounding Output bounding
Summarizing clinical Defined document types and Extractive output only; every
documentation structured fields assertion linked to its source
Applying a published Retrieval restricted to a curated, Closed set of permitted
guideline versioned corpus recommendation categories
Drafting imaging Admissibility gate on modality, Closed finding vocabulary; report
reports protocol, and quality; pre- template; structured claims
specified prompts
Conversational triage Topic scope with refusal; Closed set of dispositions;
structured symptom intake mandatory escalation rules
Analysis of wearable Defined sensors and sampling Closed set of detectable events
signals rates; signal-quality gate
C2. Structure 2rst and text 2rst: drafting a knee MRI report
Structure first. The function emits a structured list of findings drawn from an allow-listed vocabulary —
for each, the entity (for example, anterior cruciate ligament tear, meniscal tear, bone marrow edema,
joint effusion), its location, and its certainty — and the narrative report is rendered from that list.
Evaluation scores the structured list directly against the reference standard. No extraction is needed.
Text first. The function emits a free-text report. A claim extractor produces the list of findings from the
text. Before that list can be scored, the extractor is itself validated against findings extracted by
radiologists from a held-out sample of reports, for both accuracy and coverage, including coverage of
findings expressed in unanticipated phrasing.
C3. Measuring claim-level reproducibility
Method. Each case is evaluated twenty times under the deployed configuration, and again under
clinically irrelevant perturbation — the clinical history reordered, the indication paraphrased, an
irrelevant sentence added. For each finding, the proportion of runs in which the function makes the
same claim is recorded.
Illustration. A claim that an anterior cruciate ligament tear is present, made in twenty of twenty runs
and unchanged under perturbation, is reproducible. A claim of a possible meniscal tear, made in eleven
of twenty runs, is not: the function cannot be said to make that claim reproducibly, whatever its
accuracy on a single run.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 41 of 45
Docket No. FDA-2026-N-7874 · Contents
Reporting. Claim-level agreement is reported per finding and stratified by clinical consequence, so that
variability in a claim concerning a time-critical finding is visible rather than averaged away.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 42 of 45
Docket No. FDA-2026-N-7874 · Contents
References
Preprint and non-peer-reviewed sources are identified as such.
1. U.S. Food and Drug Administration. Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback. Docket No. FDA-2026-N-7874. August 2026.
https://www.fda.gov/media/194242/download
2. International Medical Device Regulators Forum, Software as a Medical Device Working Group. “Software as a Medical
Device”: Possible Framework for Risk Categorization and Corresponding Considerations. IMDRF/SaMD
WG/N12FINAL:2014. September 2014.
3. U.S. Food and Drug Administration. Software as a Medical Device (SAMD): Clinical Evaluation. Guidance for Industry and
Food and Drug Administration Staff. December 2017.
4. U.S. Food and Drug Administration. Assessing the Credibility of Computational Modeling and Simulation in Medical
Device Submissions. Guidance for Industry and Food and Drug Administration Staff. November 2023.
5. U.S. Food and Drug Administration. Considerations for the Use of Artificial Intelligence To Support Regulatory DecisionMaking for Drug and Biological Products. Draft Guidance for Industry. January 2025.
6. MacMahon H, Naidich DP, Goo JM, et al. Guidelines for Management of Incidental Pulmonary Nodules Detected on CT
Images: From the Fleischner Society 2017. Radiology. 2017;284(1):228–243. doi:10.1148/radiol.2017161659
7. Guyatt G, Rennie D, Meade MO, Cook DJ, eds. Users’ Guides to the Medical Literature: A Manual for Evidence-Based
Clinical Practice. 3rd ed. New York: McGraw-Hill Education; 2015.
8. U.S. Food and Drug Administration. Multiple Function Device Products: Policy and Considerations. Guidance for Industry
and Food and Drug Administration Staff. July 2020.
9. Rashkin H, Nikolaev V, Lamm M, et al. Measuring Attribution in Natural Language Generation Models. Computational
Linguistics. 2023;49(4):777–840. doi:10.1162/coli_a_00486
10. Maynez J, Narayan S, Bohnet B, McDonald R. On Faithfulness and Factuality in Abstractive Summarization.
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020:1906–1919.
11. Ji Z, Lee N, Frieske R, et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys.
2023;55(12):Article 248. doi:10.1145/3571730
12. Magesh V, Surani F, Dahl M, Suzgun M, Manning CD, Ho DE. Hallucination-Free? Assessing the Reliability of Leading AI
Legal Research Tools. Journal of Empirical Legal Studies. 2025;22:216–242. doi:10.1111/jels.12413
13. Gao T, Yen H, Yu J, Chen D. Enabling Large Language Models to Generate Text with Citations. Proceedings of EMNLP.
2023:6465–6488. doi:10.18653/v1/2023.emnlp-main.398
14. U.S. Food and Drug Administration. Clinical Decision Support Software. Guidance for Industry and Food and Drug
Administration Staff. January 2026.
15. Dratsch T, Chen X, Rezazade Mehrizi M, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence
BI-RADS Suggestions on Reader Performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176
16. Jabbour S, Fouhey D, Shepard S, et al. Measuring the Impact of AI in the Diagnosis of Hospitalized Patients: A
Randomized Clinical Vignette Survey Study. JAMA. 2023;330(23):2275–2284. doi:10.1001/jama.2023.22295
17. Gong EJ, Bang CS, Lee JJ, Baik GH. Knowledge-Practice Performance Gap in Clinical Large Language Models:
Systematic Review of 39 Benchmarks. Journal of Medical Internet Research. 2025;27:e84120. doi:10.2196/84120
18. Vishwanath K, Alyakin A, Alber DA, Lee JV, Kondziolka D, Oermann EK. Medical large language models are easily
distracted. arXiv:2504.01201. 2025. [Preprint]
19. Zakka C, Shad R, Chaurasia A, et al. Almanac — Retrieval-Augmented Language Models for Clinical Medicine. NEJM AI.
2024;1(2):AIoa2300068. doi:10.1056/AIoa2300068
20. Pan L, Zhang Y, Yang Q, Li T, Chen Z. Long-tailed Medical Diagnosis with Relation-aware Representation Learning and
Iterative Classifier Calibration. Computers in Biology and Medicine. 2025;188:109772.
doi:10.1016/j.compbiomed.2025.109772
21. Geng S, Josifoski M, Peyrard M, West R. Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning.
Proceedings of EMNLP. 2023:10932–10952. doi:10.18653/v1/2023.emnlp-main.674
22. Scholak T, Schucher N, Bahdanau D. PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from
Language Models. Proceedings of EMNLP. 2021:9895–9901. doi:10.18653/v1/2021.emnlp-main.779
23. Willard BT, Louf R. Efficient Guided Generation for Large Language Models. arXiv:2307.09702. 2023. [Preprint]
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 43 of 45
Docket No. FDA-2026-N-7874 · Contents
24. Inan H, Upasani K, Chi J, et al. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.
arXiv:2312.06674. 2023. [Technical report]
25. Rebedea T, Dinu R, Sreedhar M, Parisien C, Cohen J. NeMo Guardrails: A Toolkit for Controllable and Safe LLM
Applications with Programmable Rails. Proceedings of EMNLP: System Demonstrations. 2023:431–445.
doi:10.18653/v1/2023.emnlp-demo.40
26. Park K, Wang J, Berg-Kirkpatrick T, Polikarpova N, D’Antoni L. Grammar-Aligned Decoding. Advances in Neural
Information Processing Systems. 2024. arXiv:2405.21047
27. Geng S, Cooper H, Moskal M, et al. JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language
Models. arXiv:2501.10868. 2025. [Preprint]
28. Zhang S, Zhao J, Dong H, et al. When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs
with Structured Output. arXiv:2503.24191v3. 2026. [Preprint; to appear at ACM CCS 2026]
29. 21 CFR 860.7, Determination of safety and effectiveness, including paragraph (c)(2), definition of valid scientific
evidence.
30. U.S. Food and Drug Administration. Technical Performance Assessment of Quantitative Imaging in Radiological Device
Premarket Submissions. Guidance for Industry and Food and Drug Administration Staff. June 2022.
https://www.fda.gov/media/123271/download
31. U.S. Food and Drug Administration. De Novo Classification Request for IDx-DR (DEN180001): Decision Summary. 21
CFR 886.1100, Retinal diagnostic software device; Class II; product code PIB. April 2018.
https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf
32. Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC. Pivotal trial of an autonomous AI-based diagnostic system for
detection of diabetic retinopathy in primary care offices. npj Digital Medicine. 2018;1:39. doi:10.1038/s41746-0180040-6
33. U.S. Food and Drug Administration. Executive Summary for the Digital Health Advisory Committee Meeting: Total
Product Lifecycle Considerations for Generative AI-Enabled Devices. November 2024.
https://www.fda.gov/media/182871/download
34. U.S. Food and Drug Administration. Consideration of Uncertainty in Making Benefit-Risk Determinations in Medical
Device Premarket Approvals, De Novo Classifications, and Humanitarian Device Exemptions. Guidance for Industry
and Food and Drug Administration Staff. August 2019.
35. He H, Thinking Machines Lab. Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism.
September 2025. [Industry technical report; not peer-reviewed] https://thinkingmachines.ai/blog/defeatingnondeterminism-in-llm-inference/
36. Vara. Vara Receives World-First CE Certification for Autonomous AI in Breast Cancer Screening. Business Wire,
September 2, 2026; as reported in: Vara receives CE mark for first autonomous breast cancer screening. MedTech
Dive. September 2026. https://www.medtechdive.com/news/vara-receives-ce-mark-for-first-autonomous-breastcancer-screening/829696/
37. U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan
for Artificial Intelligence-Enabled Device Software Functions. Guidance for Industry and Food and Drug
Administration Staff. August 2025.
38. Kadavath S, Conerly T, Askell A, et al. Language Models (Mostly) Know What They Know. arXiv:2207.05221. 2022.
[Preprint]
39. Eisemann N, Bunk S, Mukama T, et al. Nationwide Real-World Implementation of AI for Cancer Detection in PopulationBased Mammography Screening. Nature Medicine. 2025;31(3):917–924. doi:10.1038/s41591-024-03408-6
40. Zhang Y, Chen L, Zhao W, et al. Vision-Language Models vs Autonomous AI Agents for Anterior Capsular Radial Folds:
A Diagnostic Study. medRxiv. 2026. doi:10.64898/2026.01.15.26344200. [Preprint]
41. Avakian A. Vision–Language Models as an Integrative Layer for Clinical Artificial Intelligence in Radiology: A SystemsLevel Perspective. Clinical Imaging. 2026;136:110841. doi:10.1016/j.clinimag.2026.110841
42. Mayumu N, Khan Z, Stephens M, Mukala P, Oroumchian F. RVLM: Recursive Vision-Language Models with Adaptive
Depth. arXiv:2603.24224. 2026. [Preprint]
43. U.S. Food and Drug Administration. The 510(k) Program: Evaluating Substantial Equivalence in Premarket Notifications
[510(k)]. Guidance for Industry and Food and Drug Administration Staff. July 2014.
44. Oakden-Rayner L, Dunnmon J, Carneiro G, Ré C. Hidden Stratification Causes Clinically Meaningful Failures in Machine
Learning for Medical Imaging. Proceedings of the ACM Conference on Health, Inference, and Learning. 2020:151–
159. doi:10.1145/3368555.3384468
45. D’Amour A, Heller K, Moldovan D, et al. Underspecification Presents Challenges for Credibility in Modern Machine
Learning. Journal of Machine Learning Research. 2022;23(226):1–61.
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 44 of 45
Docket No. FDA-2026-N-7874 · Contents
46. Windecker D, Baj G, Shiri I, et al. Generalizability of FDA-Approved AI-Enabled Medical Devices for Clinical Use. JAMA
Network Open. 2025;8(4):e258052. doi:10.1001/jamanetworkopen.2025.8052
47. Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable Generalization Performance of a Deep
Learning Model to Detect Pneumonia in Chest Radiographs: A Cross-Sectional Study. PLOS Medicine.
2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683
48. Badano A, Graff CG, Badal A, et al. Evaluation of Digital Breast Tomosynthesis as Replacement of Full-Field Digital
Mammography Using an In Silico Imaging Trial. JAMA Network Open. 2018;1(7):e185474.
doi:10.1001/jamanetworkopen.2018.5474
49. U.S. Food and Drug Administration. Medical Device Development Tools (MDDT). Program webpage. Accessed
September 2026. https://www.fda.gov/medical-devices/medical-device-development-tools-mddt
50. Cognita Imaging. Cognita Imaging Receives $1.29 Million FDA Contract to Test New Approach to Evaluating Generative
AI in Radiology. Business Wire. September 16, 2026.
https://www.businesswire.com/news/home/20260916329852/en/
51. National Institute of Standards and Technology. Face Technology Evaluations — FRTE/FATE. Program webpage.
Accessed September 2026. https://www.nist.gov/programs-projects/face-technology-evaluations-frtefate
52. U.S. Food and Drug Administration. Qualification of Medical Device Development Tools. Guidance for Industry, Tool
Developers, and Food and Drug Administration Staff. July 2023. https://www.fda.gov/media/87134/download
53. Wen PY, Macdonald DR, Reardon DA, et al. Updated Response Assessment Criteria for High-Grade Gliomas: Response
Assessment in Neuro-Oncology Working Group. Journal of Clinical Oncology. 2010;28(11):1963–1972.
doi:10.1200/JCO.2009.26.3541
54. Wingerchuk DM, Banwell B, Bennett JL, et al. International Consensus Diagnostic Criteria for Neuromyelitis Optica
Spectrum Disorders. Neurology. 2015;85(2):177–189. doi:10.1212/WNL.0000000000001729
Comments of UltraAI · Generative AI-Enabled Medical Devices Page 45 of 45