TechInHSR LLC (Tamiko Eto)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
Please see attached for full comments. Summary below: We appreciate FDA’s continued efforts to develop a risk-based framework for generative AI-enabled medical devices. While we support exploring a competency-based approach, evaluation frameworks must be firmly anchored in real-world workflows, specified intended use, and fundamental human rights, including human dignity, patient privacy, non-discrimination, and the right to safety and a remedy, aligning with the OECD AI Principles, Universal Guidelines for AI (UGAI) , and Center for AI and Digital Policy (CAIDP) advocacy for AI risk thresholds . Specifically, effective premarket oversight requires establishing a valid clinical association before analytical benchmarking. Post market monitoring must serve to verify premarket performance rather than substitute for deficient premarket evidence. Crucially, unless we ensure safe, trustworthy AI in healthcare supported by clear guardrails and defined lines of accountability so that clinicians and hospitals are not left carrying sole liability for faulty AI, healthcare providers will lack the confidence to adopt these technologies, undermining sustainable investment in medical AI innovation.
Attachment
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 1
Radiological Health (CDRH), under the FDA
Responses to Discussion Paper:
Considerations for the Regulation of Generative AI-Enabled Medical Devices
Docket No. FDA-2026-N-7874
September 2026
Summary
We appreciate FDA’s continued efforts to develop a risk-based framework for generative AIenabled medical devices. While we support exploring a competency-based approach, evaluation
frameworks must be firmly anchored in real-world workflows, specified intended use, and
fundamental human rights, including human dignity, patient privacy, non-discrimination,
and the right to safety and a remedy, aligning with the OECD AI Principles, Universal
Guidelines for AI (UGAI)1, and Center for AI and Digital Policy (CAIDP) advocacy for AI risk
thresholds2. Specifically, effective premarket oversight requires establishing a valid clinical
association3 before analytical benchmarking. Post market monitoring must serve to verify
premarket performance rather than substitute for deficient premarket evidence. Crucially, unless
we ensure safe, trustworthy AI in healthcare4 supported by clear guardrails and defined lines of
accountability so that clinicians and hospitals are not left carrying sole liability for faulty AI,
healthcare providers will lack the confidence to adopt these technologies, undermining
sustainable investment in medical AI innovation.
Recent Agency Actions Regarding AI in Healthcare
We applaud the FDA’s recent efforts in updating AI oversight in healthcare to support
innovation while protecting patient safety, reflecting the OECD AI Principles and Universal
Guidelines for AI (UGAI) on robust, safe, and transparent AI. While FDA pilot programs and
regulatory sandboxes offer promising pathways for iterative oversight, any operational
adjustments or regulatory flexibility must be strictly contingent upon establishing robust
premarket clinical evidence, independent verification, and explicit safeguards for patient safety
and human rights. The path to fully aligning innovation with safe, effective use remains an
ongoing effort.
Below are our responses and recommendations:
Key Recommendations
We submit these comments in support of the FDA’s commitment to advance the responsible
integration of generative AI-enabled medical devices. These comments address regulatory
challenges, opportunities for reducing burden while strengthening evidentiary quality, and
https://oecd.ai/en/catalogue/tools/universal-guidelines-for-artificial-intelligence
https://www.caidp.org/reports/caidp-index-2025/&opi=89978449&psig=AOvVaw2B7f62usp5tF0JtwN8ec3&ust=1790080935127000
https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation
https://www.pharmalive.com/the-biggest-barrier-to-ai-adoption-in-healthcare-isnt-regulation-its-trust/
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 2
Radiological Health (CDRH), under the FDA
specific guidance needs, anchored in the OECD AI Principles, Universal Guidelines for AI
(UGAI), and CAIDP statements on human-centric AI governance.5
1) Risk assessments must be based on the device’s intended use, clinical task, and realworld workflow.
2) Clinical association must be established before conducting any analytical benchmarking
or determining premarket evidence levels.
3) Human-AI interactions and agentic workflows must be evaluated holistically, including
multi-turn conversational trajectories.
4) Require risk-proportionate clinical validation, ranging from retrospective reviews to
prospective clinical studies, for all devices prior to market entry.
5) Enforce mandatory transparency and disclosures from third-party foundation model
developers regarding training data, safety limitations, and updates.
6) Utilize independent third parties to verify benchmarking, fairness assessments, and
performance for medium- and high-risk devices to mitigate conflicts of interest.
7) Design postmarket monitoring to verify premarket performance claims and detect drift,
rather than to substitute for inadequate premarket evidence.
RESPONSES to: Section IV: Considerations for the Assessment of Risk for GenAIEnabled Devices - Discussion Questions
1. Does the two-axis risk framework, organized around device activity and the
consequence of relying on an incorrect output, appropriately capture the
dimensions most relevant to the risk of a GenAI-enabled software function?
○ No. We recommend adding intended use and scope as explicit considerations. In
accordance with OECD and UGAI Principles regarding risk management, as well
as CAIDP advocacy on AI risk thresholds6 for advanced AI systems, a device’s
risk cannot be assessed without first understanding what it is intended to do, for
whom, and in what clinical setting.
○ We also recommend considering reversibility, human intervention, time pressure,
and traceability. An action that can be readily reversed and reviewed by an HCP
presents a different risk from an autonomous action with serious and irreversible
consequences.
2. What characteristics of an output-such as its wording, specificity, personalization,
or context-could be considered as modifiers of the risk associated with an
informational function, after accounting for the device’s overall functionality and
intended use? What additional information would provide manufacturers with
sufficient clarity and predictability around risk assessment for informational
Michael Karanicolas, Artificial Intelligence and Regulatory Enforcement, Report for the Administrative Conference of the
United States, Dec. 9, 2024, pg. 10, https://www.acus.gov/sites/default/files/documents/AI-Reg-Enforcement-Final-Report2024.12.09.pdf
https://www.techpolicy.press/the-ai-red-line-challenge/
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 3
Radiological Health (CDRH), under the FDA
functions while recognizing that directiveness may exist along a continuum rather
than as a binary distinction?
○ We recommend that the FDA evaluate directiveness based on real-world clinical
workflow rather than user interface language or disclaimers, such as "talk to your
doctor," which should not be treated as risk mitigations. Directiveness should be
assessed on a continuum by evaluating the frequency with which clinicians
override or ignore device outputs in practice, where higher override rates reflect
lower directiveness and rare overrides indicate stronger guidance of decisionmaking.
○ Risk assessments account for contextual factors such as clinical reviewability,
action reversibility, explanation transparency, outcome severity, patient
population vulnerability, and representative dataset testing.
3. What device characteristics, output features, or safeguards might mitigate risks that
could arise when a user lacks the domain knowledge to independently evaluate an
output, without unnecessarily underestimating patient capability?
○ We recommend that FDA treat patient-facing informational tools as presenting
higher risk when addressing high-stakes health topics or decisions, such as triage
for emergency care. Risk determination should be anchored in the specific clinical
task, the potential impact of output errors, and the level of domain knowledge
required to evaluate system recommendations, rather than user classification
(patient vs provider) alone.
○ Furthermore, we recommend that regulatory risk assessments avoid using user
role as a sole risk determination factor and instead mandate risk-proportionate
safeguards for high-risk applications, including direct emergency care advisories,
confidence indicators for uncertain outputs, and required human clinical review
prior to high-stakes action.
4. Under what circumstances, if any, might risk be affected when an HCP who lacks
the relevant clinical specialist knowledge receives information that falls within a
specialist area of practice? What device characteristics or safeguards might mitigate
such risks?
○ We recommend focusing on intended use and the qualifications required for safe
use, rather than a generalist-versus-specialist distinction.
○ A generalist may safely use a device within its intended scope, while a specialist
may not be appropriately qualified to use a device outside their area of expertise.
The intended use should therefore identify the relevant user and clinical setting.
Where necessary, institutional controls and workflow restrictions can help prevent
use outside that scope.
5. For multi-turn conversational GenAI-enabled devices that may migrate from
providing “non-directive” information to “action-directing” information over the
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 4
Radiological Health (CDRH), under the FDA
course of an exchange, how should risk be assessed across realistic conversational
trajectories? How could the intended use of such a device be characterized when its
behavior is emergent across a conversation?
○ Risk assessments should evaluate entire conversations rather than just single
responses, as AI behavior can change as a chat progresses. Manufacturers should
use multi-turn simulations to test whether the AI recognizes when a conversation
has drifted outside its scope or into a risky or “directive” area, ensuring that it
properly escalates to a human clinician rather than continuing autonomously.
○ The "intended use" statement must clearly define the AI's conversation
boundaries, such as limits on the number of turns or the types of medical advice
allowed. This prevents "scope creep" where an informational tool inadvertently
becomes a diagnostic or treatment-directing tool over a long, open-ended
exchange.
6. For GenAI-enabled care escalation functions, CDRH is considering how the
evaluation may account for both under-escalation and over-escalation. How could
manufacturers characterize and weigh these two directions of error, given that they
may not be commensurable and that acceptable trade-offs may vary by clinical
context?
○ We recommend benchmark testing that measures missed escalations and
unnecessary escalations. Missing a needed escalation can delay care and cause
serious harm. At the same time, unnecessary escalation may waste resources.
Both errors should be measured separately rather than combined into a single
performance measure. The appropriate balance will depend on the clinical context
and the consequences of each error.
RESPONSES to: Section V: Competency Based Approach for Premarket Evaluation of
GenAI-Enabled Devices - Discussion Questions
7. Is the competency-based approach described above, i.e., device benchmarking
followed by clinical confirmation, a useful and appropriate framework for
evaluating GenAI-enabled devices?
○ Before device benchmarking can occur, it must first be established what the
device is intended to do and how it is actually used. This means seeking clinical
confirmation and asking about clinical practice. Such evidence is more efficiently
collected during a research study to understand how the output would be used,
who has authority to override it, what happens when clinicians do not follow the
recommendations, and what workflow changes it enables. Research applications
test these factors to determine whether the research protections were sufficient,
insufficient, or excessive. The research study also serves to confirm that the
clinical outcome the device predicts actually correlates with the device’s outputs
in the intended-use population (clinical association). This approach has been
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 5
Radiological Health (CDRH), under the FDA
taken for decades under 21 CFR 812 and 820/ISO 13485. Only after these steps
are taken can we move on to risk classification for postmarket deployment.
○ We support the proposed combination of benchmarking and clinical confirmation,
provided that three things are included:
i. both must be proportionate to the device’s intended use and risk;
benchmarking outside the intended use is not appropriate evidence;
ii. before clinical deployment, institutions must demonstrate the availability
of regulatory expertise on staff, documented device-determination
processes under 21 CFR 860 and 820/ISO 13485, quality management
systems that include device oversight separate from the IRB, a monitoring
protocol, access controls to document out-of-scope use, and staff training
on intended use and device limitations; and
iii. premarket evidence should include benchmarking plus a clinical
confirmation method, such as shadow deployment or simulation.
Postmarket monitoring may supplement, but not substitute for, these
premarket requirements.
8. How might the two-axis risk framework described in Section IV be considered
within the competency-based approach described above to help determine the level
of evidence needed for a premarket submission?
○ The two-axis risk framework cannot and should not serve as a standalone risk
assessment. This framework is meaningful only after a valid clinical association
has been established. If the framework is applied without first establishing clinical
association, the device is being plotted on a matrix without an adequate
understanding of what the device actually is or what it is supposed to do. Only
after identifying the clinical problem the device addresses, the clinical association
it relies on, and how clinicians actually use it can the FDA determine how
rigorous the premarket evidence needs to be.
○ All devices must require both analytical validation and clinical validation. What
scales with risk is the rigor of the clinical confirmation: the design of the
validation study, the independence of the clinician validation, and whether
validation was retrospective or prospective.
9. Would a benchmarking structure such as the one described above be likely to
provide adequate evidence of clinical knowledge, analytic capabilities, safety
behavior, communication, and generalizability to support a reasonable assurance of
safety and effectiveness? Are there elements that are missing, redundant, or
inappropriately categorized? Are there externally developed standards that could
be leveraged?
○ The FDA and IMDRF have relied on the Clinical Evaluation guidance for SaMD7
for over a decade, and this guidance has worked well to meet and address
https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 6
Radiological Health (CDRH), under the FDA
established global standards. This question is focused on the wrong problem. The
real question should be whether the benchmarking measurement actually predicts
real-world performance. We must first ask whether there is a scientifically
established relationship between what the device measures and the targeted
clinical outcomes. This clinical association must first be established and grounded
in clinical evidence before benchmarking, which could be equated to “analytical
validation,” can occur.
○ A benchmark can measure exactly the right thing, such as clinical knowledge or
safety-critical recognition, and still be meaningless if it does not reflect actual
deployment. If the benchmark uses cases that are simpler than real cases, from a
different population, or formatted differently from what clinicians see, then
passing the benchmark says nothing about whether the device will work safely
when actually deployed (Choi et al. 2023, Cassio et al. 2023, and Ramwala et al.
2024). A benchmark can be validated but have no clinical basis. A benchmark can
also be perfectly designed but fail to predict real-world performance if it does not
reflect actual deployment conditions. We strongly urge the FDA to ensure that
manufacturers prospectively conduct clinical-association analyses and specify
acceptable performance before running the benchmark. The FDA should bring
back the Clinical Evaluation guidance that was recently archived, as it clarifies
how to apply this long-established framework for AI/ML and generative AI in
medical devices.
10. CDRH recognizes that publicly available benchmarking assets may be subject to
data contamination, saturation, and limited real-world representativeness. How
should a sponsor establish that performance on a given benchmark predicts safe
and effective real-world behavior for the device’s intended use? What evidence
should support the construct validity of a benchmark used to gate device evaluation,
and what role should sponsor-developed benchmarks play given potential concerns
around independence and optimization to the test?
○ Proper benchmarking asks, simply, whether the device performs well when
deployed clinically. To establish this, the manufacturer must compare benchmark
cases to real cases and track real-world performance against benchmark
predictions. The manufacturer must be transparent about the synthetic data used
and must use independent expert review for sponsor-developed benchmarks. The
main problem with sponsor-developed benchmarks is the conflict of interest,
which has historically been heavily scrutinized by federal agencies, including the
FDA. The manufacturer will naturally create benchmarks that play to the device’s
strengths, avoid exposing its weaknesses, and sometimes do so because that is
easier, faster, and cheaper than real-world performance testing. Sponsors may
exclude edge cases, simplify scenarios to make them more like training data, or
weight performance dimensions in ways that favor how the device behaves. To
mitigate this “optimization bias,” we strongly urge the FDA to continue requiring
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 7
Radiological Health (CDRH), under the FDA
independent experts—not employees of the manufacturer and not compensated
based on whether the device passes—to assess whether the benchmark is
genuinely representative of real deployment or has been optimized in favor of the
manufacturer. Independent expert reviewers may assess whether the cases are
realistic, whether they are drawn from the actual distribution of cases clinicians
are likely to see, and whether the benchmark exposes weaknesses that should be
made transparent to the clinicians using it. They should ask whether the scoring
rubric reflects what actually matters clinically or weighs dimensions in ways that
advantage the device. Benchmark cases should be compared with real cases to
verify a distribution match. After deployment, the FDA should require tracking of
whether real-world performance matches what the benchmark predicted. If realworld performance diverges significantly from benchmark predictions, the
benchmark was not valid; it was likely measuring something other than real-world
capability and needs to be redesigned. We strongly urge the FDA to require
manufacturers to be transparent about the synthetic data used, document how
these synthetic cases were created, identify what distribution they follow, and
explain whether that distribution is representative of real cases. The danger with
synthetic data is that it often reflects the biases of the model that generated it.
Therefore, if a benchmark consists of synthetic cases, the benchmark may not
catch the model’s real-world failures.
11. CDRH is considering that clinical confirmation for a GenAI-enabled device might
not require a prospective clinical study in every case, and has described above a
range of approaches of increasing rigor and patient exposure. How might a sponsor
select and justify a confirmation approach tailored to a device’s intended use and
proportionate to the device’s risk profile? Are there device types or risk profiles for
which one or more of these approaches would be insufficient or inappropriate? How
might the anticipated distribution of real-world inputs be taken into consideration?
Are there other methods of clinical confirmation that might help inform the
evaluation of GenAI-enabled devices?
○ Only after a device determination is made and the device risk level is identified
based on intended-use statements can the sponsor determine which type of
clinical confirmation is appropriate. Clinical confirmation must be required for all
devices because it validates that the device actually performs its intended function
in real deployment. The risk level determines how rigorous that confirmation
must be. For low-risk devices, clinical confirmation is required, but lighter
methods are acceptable, such as a retrospective review of real patient records
checking whether the device’s recommendation aligns with what clinicians
actually did, or whether clinical validation samples are sufficient. In this case,
prospective clinical trials would not likely be needed, but evidence that the device
performs as intended when clinicians actually use it is required. Medium-risk
devices require clinical confirmation with more rigor. Retrospective review of
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 8
Radiological Health (CDRH), under the FDA
real patient records, shadow deployment or simulation, or prospective clinician
validation samples are appropriate in these cases. One of these methods should be
included to demonstrate that the device performs safely when integrated into real
clinical decision-making. High-risk devices require the most rigorous clinical
confirmation. Retrospective review or shadow deployment is insufficient. We
recommend prospective clinical evidence, ideally a clinical study, showing that
the device performs safely and effectively when deployed with real patients,
clinicians, and real clinical complexity, and demonstrating that the device’s
autonomous or action-directing recommendations actually lead to safe and
effective outcomes. Whichever approach is used, clinical confirmation must
include the populations for which the device is intended. If the device is intended
for use across different age groups, racial groups, or disease severity levels, then
the clinical confirmation must test it in those actual populations. A sponsor cannot
skip population-specific validation by claiming that the sample was too small.
12. How should sponsors achieve statistically meaningful performance measurement for
GenAI-enabled devices? Where synthetically generated inputs supplement real
patient data, how should sponsors account for differences between the synthetic and
real-world distributions when estimating performance, and under what conditions,
if any, is it appropriate to combine benchmarking evidence and clinical
confirmation evidence to support a single performance estimate?
○ Traditional medical devices allow sponsors to calculate sample size up front.
However, with generative AI devices that produce open-ended outputs, it may not
be possible to pre-calculate sample size because there is no fixed endpoint.
Therefore, GenAI devices need a different approach. Before asking “how much
evidence is enough,” device determinations must first establish whether a clinical
association actually exists between what the device measures and what it claims
to predict. This prerequisite cannot be skipped. Device determinations must first
answer whether there is a real clinical association. In other words, does the device
actually measure what predicts the outcome in the intended-use population, or is it
measuring something else? Only once the clinical association is established can
the sponsor ask how much evidence is needed to demonstrate device-performance
reliability. Once the clinical association is established, the quality and
representativeness of the data must be evaluated. The validation population used
to establish adequate performance must be representative of the intended-use
population. If it is not, poor performance signals deployment outside the validated
scope, and no amount of data will be a sufficient sample size. Only after that is
established should the sponsor define what acceptable performance means. For
example, if clinician validators independently agree that the device’s
recommendation is appropriate in 80% of cases, then evidence should be collected
until that statistical confidence is achieved. At that point, it may be necessary to
continue collecting data until the confidence interval is met, keeping in mind that
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 9
Radiological Health (CDRH), under the FDA
data representativeness is critical. Reporting must be transparent, with point
estimates, confidence intervals, sample size, and population characteristics. But
again, if clinical association has not been established, no amount of subsequent
data collection will create one. Preliminary evidence establishing that association
is needed before advancing to benchmarking and statistical validation.
13. For which clinical domains, device functions, or subpopulations is synthetic data
particularly well-suited, or particularly inadequate, as a supplement to real-world
evidence? What safeguards would mitigate the risk that synthetic data generated by
models of the same class as the device under evaluation reproduces the very
performance gaps the evaluation is intended to detect, particularly for
underrepresented subgroups?
○ Synthetic data can help with some benchmarking needs, but it has a critical
limitation: it reproduces the biases of the model that generated it. Synthetic data
may help when, for example, a sponsor wants to evaluate variation in certain
conditions. For instance, a sponsor may have a case involving a diabetic patient
with pneumonia and may want to generate variations involving a diabetic patient
with pneumonia and renal impairment, pneumonia while taking three
medications, and so forth, to test whether the device’s output changes
appropriately with different presentations. Another circumstance in which
synthetic data may be useful is when a sponsor needs more cases of an
uncomplicated UTI to test whether the device recommends appropriately. Such
synthetic cases may supplement real cases, but they would not be good candidates
for validating safety and effectiveness. Stress testing is an excellent use case for
synthetic data. A sponsor may be able to test how the device handles unusual
inputs, such as extremely long conversations, unusual formatting, or conflicting
information.
○ However, synthetic data cannot and should not replace real data. If the device’s
training data lacked a representative dataset—for example, if it lacked diversity in
any given population—synthetic data generated from the same model will
reproduce that same lack of diversity. Sponsors cannot synthetic-data their way
out of bias in the training data. They need real data from the populations of
concern. If the concern is underrepresentation, for example, or whether the device
works for all subgroups, those questions cannot be answered with synthetic data.
For high-risk devices, we strongly recommend a fairness assessment using real
data as a condition of approval. A fairness assessment evaluates whether the
device performs equally well across different patient populations or whether
device performance diverges in ways that could harm specific groups. This
question can only be answered with real data. More importantly, postmarket
monitoring could verify that the fairness originally observed premarket holds up
in actual deployment.
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 10
Radiological Health (CDRH), under the FDA
14. For open-ended device outputs, how should performance comparators and
acceptance criteria be selected? When a panel of qualified clinicians serves as the
comparator, how should the applicable standard (for example, the standard of care
versus the performance of a median clinician in practice) be defined and justified?
Should generalist or specialist physicians be used as a performance standard? When
should human-AI team performance, rather than the device operating alone, serve
as the basis for evaluation?
○ For open-ended outputs, there must be a clear standard for comparing the device’s
performance against something. That standard depends on what is available and
what the manufacturer claims the device can do. A true reference standard, or
ground truth, is the best option. For some domains, this exists. For example, if the
device recommends a diagnosis, pathology results are ground truth. If the device
recommends treatment following published guidelines, those guidelines are
ground truth. However, when ground truth does not exist, clinician consensus
must be used to compare the AI output against clinician performance. To do this,
multiple independent clinicians should review the same cases and make their own
recommendations. The sponsor should then measure whether the device performs
as well as—or better than—the median clinician. Not all clinicians agree on
everything, but if the device can meet median clinician performance, that is a
reasonable standard. For high-stakes decisions, we recommend comparison
against specialist clinician performance and not median clinician performance.
For example, a cardiologist using a cardiology device should be compared with
board-certified cardiologists, not generalists. For novel areas where neither
ground truth nor clinical consensus exists, structured expert opinion should be
used, in which qualified experts independently assess whether the device’s
outputs are appropriate using a validated scoring rubric. The comparator should
be prespecified up front. Prespecifications should include the qualifications of the
comparator and assurance processes that confirm independence from the
manufacturer. The scoring rubric should be published beforehand, not after the
results are known. Appropriate output must be defined ahead of time.
Performance thresholds must be established before testing, and the reference
population must be specified in advance. For example, is the device intended for
generalists, specialists, adult populations, or pediatric populations? Is it intended
for simple or complex cases? This is where the intended-use statement plays a
critical role.
15. Are there ways in which performance might be assessed relative to the care,
technology, or course of action likely to occur in the absence of the device, rather
than to the comparators described in this section? What approaches might be used
to identify and justify a comparator such as unaided clinical judgment, delayed
specialist review, or no intervention?
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 11
Radiological Health (CDRH), under the FDA
○ When the comparators described above are not feasible or appropriate, then other
valid comparisons may exist depending on the device’s purpose. For example, the
status quo (how clinicians currently work without the device) is the best approach.
For some devices, the evidence can show that clinicians using the device perform
better than clinicians not using the device. This requires studying the same
clinicians or equivalent clinicians with and without the device. Another approach
may be delaying specialist review under a controlled IRB-overseen study where
triage or routing functions are evaluated and such evidence might show that the
device correctly routes cases to the right specialist or correctly identifies which
cases need expert review, thereby reducing specialist burden or improving routing
accuracy. Another approach would be under an IRB-approved study when there is
no-intervention baseline, and there are no doctors who know yet what the best
approach is, you can show that patients with the device don’t have worse
outcomes than the patients without it. Critically, whatever alternative comparator
you choose, the comparator must be specified in advance and justified why it's
appropriate for your device’s intended use. A triaging device for an ED can use
“did we correctly route to the right department?” as a comparator and a drugselection device needs a more rigorous comparator. A device claiming to reduce
clinician bias would need evidence showing it actually does reduce bias and not
just change clinician behavior.
16. Should independent third parties be involved in some or all aspects of a
competency-based approach, including device benchmarking and clinical
confirmation? If so, in what ways might qualified, independent third-party
participation contribute to this approach, and what qualifications and independence
criteria should apply? Are there aspects of the assessment for which third-party
involvement would be impractical or inadvisable? What safeguards or program
design features would be critical to prevent third-party participation from limiting
competition or preventing innovation?
○ When a manufacturer benchmarks its own device, approval presents a conflict of
interest. This is something the FDA has evaluated closely for decades. In such
cases, the company has every reason to make the benchmark favorable.
Independent review does not eliminate bias, but it adds a check that the
benchmark is reasonable. For high-risk devices involving action-taking,
significant consequences, or autonomous action, and for medium-risk devices, the
FDA should require independent third parties for critical functions such as
benchmarking verification, clinician validation, and fairness-assessment review.
Independent review should ensure that there is no financial relationship with the
manufacturer or the foundation model developer. Reviewers must have real
expertise in the device’s clinical domain. This review should not be done by
another AI tool. For low-risk devices, third-party review may be optional, but
evidence without such review should be seen as less credible.
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 12
Radiological Health (CDRH), under the FDA
17. CDRH is exploring whether a competency-based approach could be applied to
devices that incorporate a variety of underlying model architectures, such as
multimodal vision-language models and generative or predictive world models. Are
there aspects of the described approach that might be ineffective or inapplicable to
these devices? What additional or distinct considerations might need to be
addressed?
○ For multimodal systems that combine vision and language, the existing
benchmarking framework adapts naturally. Some elements apply differently
depending on the input type. For example, quantitative analysis applies when the
device processes structured data fields, while information gathering applies when
it interprets images. An additional consideration is cross-modal consistency. For
example, if the device is presented with the same clinical case as an image and
then as a text description, does it recommend the same thing? Inconsistency
would indicate that the device is relying on superficial features rather than
understanding the underlying clinical problems. For generative AI systems that
predict trajectories or simulate patient outcomes, the framework also applies. For
example, the intended-use statement should specify what the predictions are
meant to predict, such as whether the device predicts deterioration in the next 72
hours and not an entire hospitalization. For these stochastic systems, in which
randomness is built in, reproducibility is critical. If the device generates slightly
different predictions each time it is asked the same question, that variability needs
to be understood and bounded. Clinical confirmation would validate that the
predicted trajectories actually match what happens in real patients.
RESPONSES to: Section VI: Postmarket Monitoring for GenAI-Enabled Devices -
Discussion Questions
18. CDRH is considering whether it may be appropriate to accept greater premarket
uncertainty regarding a GenAI-enabled device’s benefit-risk profile through greater
reliance on postmarket monitoring. Under what conditions might such an approach
be appropriate, and what characteristics of a monitoring program would need to be
in place to justify reduced premarket evidence? Are there device types or risk
profiles for which this approach would not be appropriate?
○ The amount of premarket evidence required depends on the device’s risk profile.
Low-risk devices, such as those that provide only informational output and have
limited consequences if wrong, can rely on benchmarking alone. Medium-risk
devices that are more action-directing and have moderate consequences would
need benchmarking in addition to one method of clinical confirmation, such as
shadow deployment or clinician validation, to show that the device actually works
when clinicians use it. High-risk devices that involve autonomous action or
significant consequences need comprehensive benchmarking plus prospective
clinical evidence from a clinical study before approval.
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 13
Radiological Health (CDRH), under the FDA
○ We highly recommend against framing postmarket monitoring as a substitute for
inadequate premarket evidence, consistent with CAIDP statements on
accountability, Universal Guidelines for AI (UGAI), and OECD principles for AI
safety throughout the lifecycle. Postmarket monitoring’s job is to verify that what
you claim premarket actually holds true in real deployment. It can catch
unexpected problems, but it cannot discover problems that premarket testing was
designed to catch and missed. If you test a device inadequately premarket, then no
amount of postmarket monitoring is going to fix that. It will only show you that
your premarket testing was inadequate and by then the device has already harmed
patients. The consequence is that premarket evidence is insufficient for a device’s
risk profile and should not be offered conditional approval pending postmarket
data. Postmarket monitoring verifies that the premarket claims hold up.
19. Please comment on the potential approaches to postmarket performance evaluation,
including periodic re-benchmarking, sample-based clinician review, and
performance degradation monitoring. What additional approaches should CDRH
consider, and how should the cadence and triggering events for reassessment be
determined?
○ For low-risk devices, annual performance evaluation may be adequate, and the
primary reporting mechanisms would be standard adverse event reporting. For
example, currently, when something goes wrong, clinicians report it. For mediumrisk devices, monitoring becomes more systematic. For example, quarterly rebenchmarking to check whether performance has degraded, annual clinician
review of sample, and annual drift monitoring to detect degradation over time. For
high-risk devices, the cadence should increase to monthly drift monitoring and
real-time tracking, quarterly clinician review, and semi-annual fairness audits to
verify that performance is consistent across relevant patient subgroups.
20. Please comment on whether the proposed approaches to postmarket monitoring can
be facilitated by machine-based supervisory agents. What considerations, including
the evaluation and reliability of the supervisory agent itself, should CDRH take into
account for such an approach?
○ Automated monitoring tools can help collect data and flag problems, but they
cannot and should not replace human clinical judgment. Appropriate uses for
machine agents may include data collection (for example automated performance
metrics evaluating accuracy on test cases and drift signal calculation; flagging
(which would include automated alerts when performance exceeds the threshold
such as a monthly accuracy dropping below a given target) and generating reports
on dashboards showing performance trends. Automated agents should not decide
whether a performance problem is clinically important, why it happened, nor
whether the device should be changed, paused, or withdrawn. If the monitoring
tool also uses Generative AI, it should be independently tested, understandable to
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 14
Radiological Health (CDRH), under the FDA
humans, and designed to send issues to people for review and auditing when
needed.
21. What roles might clinicians, healthcare institutions, professional societies,
standards-setting bodies, and other stakeholders appropriately play in postmarket
monitoring, and how can these roles be structured without diffusing manufacturer
accountability?
○ When everyone is responsible, no one is accountable. Under current enforcement,
diffused responsibility becomes diffused accountability, and each stakeholder can
claim that’s someone else’s job. Without explicit governance guardrails protecting
clinicians and health systems from unfairly absorbing sole liability for algorithmic
errors or faulty system outputs, doctors and hospitals will simply refuse to adopt
or deploy these tools. There is no sustainable incentive to invest in healthcare AI
if providers bear the ultimate risk without strict manufacturer accountability.
Instead, postmarket monitoring should assign clear roles. The manufacturer
should maintain the design history file (DHF), follow the monitoring plan, report
trends, and act on any safety signals. The deploying institution should train users,
enforce access limits, and monitor local use. FDA should set requirements,
inspect compliance, and enforce these requirements. Professional groups and
payers may support best practices, but they should not replace manufacturer or
FDA accountability.
22. CDRH envisions that a manufacturer’s premarket competency-based assessment
might serve as a baseline against which post-deployment modifications could be reevaluated. How might the extent of re-benchmarking or other evidence be scaled to
the nature and expected impact of a given modification? For example, are there
categories of change that might not significantly affect safety or effectiveness of a
GenAI-enabled device, or might be appropriately managed within a sponsor’s
quality management system versus requiring FDA premarket review and
authorization before implementation? Are there categories of changes that might be
appropriate for inclusion in a PCCP?
○ Changes after approval should be reviewed based on how much they affect the
device. Small prompt or data updates may only need focused re-testing if they
stay within a predetermined changed control plan (PCCP). New clinical uses, new
patient populations, major model changes, or large performance shifts should
require stronger review and may need FDA submission and IRB review prior to
use. This same concept carries over into Question 23 below.
23. How might PCCP concepts or other change-control approaches be adapted for
GenAI-enabled devices when the nature or scope of future modifications cannot be
fully prespecified?
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 15
Radiological Health (CDRH), under the FDA
○ For generative AI devices where, future modifications cannot be fully
prespecified, we recommend a tiered PCCP. Tier 1 may not require a submission
so long as the changes fit within prespecified boundaries, such as model updates
within +/- 2% accuracy on a held-out test set. This should include subgroup
accuracy, not just aggregate accuracy. Tier 2 should allow for expedited review
when changes fall outside Tier 1 but remain within an acceptable range, such as a
+/- 3–5% accuracy change, again considering subgroup accuracy and not just
aggregate accuracy. The manufacturer should submit a summary including rebenchmarking. A higher tier should require more intensive review when there are
significant changes, such as a 5% accuracy change or new recommendation types.
This submission would be equivalent to an original change control. Manufacturers
should monitor re-benchmarking data continuously against these tier thresholds
and escalate when those thresholds are exceeded.
24. For devices built on third-party foundation models, changes to the underlying
model may be initiated by the third-party model developer rather than the device
manufacturer. How can a manufacturer detect, evaluate, and respond to such
changes in a timely manner? What mechanisms-for example, contractual, technical,
or through a PCCP-could provide reasonable assurance that third-party developerinitiated changes do not compromise the safety or effectiveness of the device?
○ The manufacturer should be responsible for changes made by a third-party
foundation model provider. Contracts should require advance notice of model
updates, access to prior versions when rollback is needed, disclosure of important
safety changes, and audit rights for device-related performance. Foundation
model developers must notify the manufacturer within a specified period of any
model updates, including the version number, change summary, and behavioral
changes, and must guarantee that model performance will not degrade by more
than a specified percentage for a given period. This timing should depend on the
risk level of the manufacturer’s device. Manufacturers should be able to easily
revert to a prior version when needed. Third-party model developers must
proactively disclose any known changes to behavior or safety properties for
medical devices and allow transparent auditing rights to model training and
device-relevant testing. The manufacturer’s ongoing performance monitoring
should detect degradation through the PCCP described in Question 22 above. If
degradation is detected, manufacturers must assess the cause and escalate under
the PCCP. If a model change is the cause of degradation, the manufacturer must
contact the developer, assess the impact, and submit a modification or roll back to
a former version. Liability should remain with the device manufacturer when a
third-party model update harms performance and the manufacturer fails to detect
or address it. If third-party model updates cause device-performance degradation,
the manufacturer must be liable. This creates an incentive for the manufacturer to
ensure a strong contractual commitment before deployment.
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 16
Radiological Health (CDRH), under the FDA
RESPONSES to: Section VII: Other Topics
25. Would voluntary Foundation Model MAFs be practical for foundation model
developers to provide and to keep current, and what content might they include?
Given that participation would be voluntary and that model developers may have
limited incentive to disclose safety-relevant information, what would make such a
program sufficiently useful for premarket review? Are there alternative
mechanisms CDRH should consider for obtaining information about underlying
foundation models?
○ No, these programs would not work because foundation model developers have
no incentive to participate. FDA should require disclosure for foundation models
used in regulated medical devices. A voluntary system is not enough because
model developers may choose not to share information that could create legal
risk, reveal business information, lead to enforcement, or harm patients. We
recommend mandatory foundation model disclosure requirements for any model
used in FDA-regulated medical devices. Before device approval or clearance, the
manufacturer must submit the model name, version, training data categories,
known safety-critical limitations, behavioral constraints and guardrails, and
contractual commitments such as update notifications, stability, warranty, and
rollback. Some information may be confidential and only made available to the
FDA and device sponsor. If a model developer refuses to disclose these things,
the device cannot be approved. This creates a market incentive for transparency.
The manufacturer should certify that foundation model disclosures are complete
and current. This is currently the same requirement under 21 CFR 820.25/ISO
13485 which requires manufacturers to ensure suppliers provide the necessary
information.
26. Are there additional considerations that inform the premarket and postmarket
evaluation of agentic GenAI-enabled devices, beyond those applicable to nonagentic GenAI-enabled devices? How should the elevated risk associated with
autonomous multi-step action, tool use, and reduced opportunity for human review
be reflected in acceptance criteria and oversight?
○ Agentic systems are inherently higher risk than non-agentic devices and require
extra safeguards because they can take multiple steps, use outside tools and act
with less human review. In line with CAIDP's calls for explicit redlines8 and strict
oversight on autonomous systems, as well as OECD AI Principles and Universal
Guidelines for AI (UGAI)9 on human control and accountability, FDA should
require testing of multi-step behavior, safe tool use, human approval points, and
the ability to stop the system. Hospitals and clinics should keep audit trails,
review real examples of agent decisions, and monitor tool-use errors during
https://red-lines.ai/
https://oecd.ai/en/wonk/ai-governance-through-global-red-lines-can-help-prevent-unacceptable-risks
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 17
Radiological Health (CDRH), under the FDA
research and post deployment. We recommend mandatory benchmarking
elements for agentic systems. First, escalation should occur before an irreversible
action. Second, the system should refuse to go beyond its scope if a conversation
progresses. Third, benchmarking should include evaluating if the tool is used
accurately and how errors are handled. For example, does the human recognize
the wrong results? Does the system pause when a human rejects a decision? More
importantly, we cannot rely solely on benchmarking. There must be a minimum
requirement for a clinician to step in and review, with realistic multi-step
scenarios. To effectively protect humans, institutions must ensure humans are
driving critical decisions (agents would pause before irreversible action), and
there should be an audit trail showing the agent’s reasoning, and the ability for a
human to stop the agent mid-sequence. Finally, there should be regular review of
actual multi-step executions .
Conclusion
Generative AI offers transformative potential to advance healthcare delivery safely and
equitably. Realizing that potential requires a pragmatic, rigorously grounded regulatory
framework that prioritizes clinical association, workflow-based validation, and strict model
accountability over mere benchmarking, fully aligning with OECD AI Principles, Universal
Guidelines for AI (UGAI), and CAIDP recommendations for trustworthy AI governance and red
lines. Addressing these oversight needs proactively will establish a trustworthy foundation for
AI-driven medical devices, positioning the U.S. as a global leader in responsible medical
innovation.
We welcome the opportunity to contribute further to this work and are available for follow-up.
Respectfully,
Tamiko Eto, MA, CIP Vishnu Narayan, MSc, BTech
Founder & Principal Consultant, Contributing Policy Editor & Non-Resident
TechInHSR LLC Fellow at HealthTechAsia
About TechInHSR LLC
TechInHSR LLC is an AI research and regulatory compliance consulting practice specializing in AI/ML
governance, human subjects research (HSR) protection under 45 CFR 46 and 21 CFR 56. Its founder and
principal consultant, Tamiko Eto, MA, CIP, brings over twenty years of research administration
leadership experience. Ms. Eto also serves as Senior Teaching Fellow at the Center for AI and Digital
Policy (CAIDP), a global non-profit education and research center. TechInHSR’s applied frameworks,
including a three-stage risk-based AI oversight model matching oversight to project maturity, has been
adopted by numerous U.S. research hospitals & universities, and received recognition from federal
agencies.
Vishnu Narayan, a Tech policy expert, serves as a Regulatory Systems Strategist at the Medical
Comments to the Digital Health Center of Excellence (DHCoE) and Center for Devices and 2
Radiological Health (CDRH), under the FDA
Technology Association of India and as a Contributing Policy Editor and Non-Resident Fellow at
HealthTechAsia. His work focuses on the intersection of AI governance, healthcare regulation and
technology policy. He holds a Master's in Regulatory Policy and Governance from the Tata Institute of
Social Sciences and a Bachelor's in Biomedical Engineering. His work has included research and policy
contributions through CAIDP, WHO/ITU AI4H and the Commonwealth AI Consortium, with a focus on
responsible innovation, regulatory systems and healthcare resilience.