← All 95 filings

Martin Haimerl

Academia / otherAcademicFiled September 1, 20268,407 words · 1 attachmentFDA-2026-N-7874-0052
“Human oversight should be treated as a risk control whose effectiveness may require both initial and periodic evaluation, rather than as an inherently reliable safeguard.”

What they argued

RecovryAI’s one-line reading of the filing.

Supports risk-proportionate criticality stratification and competency approach; staged expansion of independence via evidence gates; clinician confirmation before irreversible actions.

Themes it raises

14 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Accordingly, safeguards should reduce criticality only where they are sufficiently defined, enforceable, verifiable, and realistic for the intended environment.”
Whether the user can judge the outputFDA Q3, Q4
“The relevant consideration should therefore be whether the intended user possesses the competence and information necessary to independently assess the particular output, rather than professional designation alone.”
Escalating too little and too muchFDA Q6
“Errors may therefore involve not only false-positive or false-negative escalation, but also incorrect escalation level, inappropriate timing, or inappropriate destination of care.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“However, I would be cautious to overemphasize the analogy to an assumption of equivalence between clinician and GenAI competency.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Average performance alone should not permit strong performance in common or low-risk situations to compensate for poor performance in a clinically critical subset.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“Prospective evidence becomes particularly important where safety depends materially on human-AI interaction, local workflow, real-time information availability, or where use of the device changes subsequent diagnostic or therapeutic decisions.”
Trading premarket certainty for postmarket monitoringFDA Q18
“In principle, increased reliance on postmarket monitoring can be appropriate, but it should not create a general trade-off whereby stronger postmarket monitoring automatically permits less premarket evidence.”
Watching the device after it shipsFDA Q19, Q20
“Postmarket monitoring should address multiple forms of drift, including data drift, model/component drift, concept or clinical drift, workflow/use drift, and oversight drift.”
Who is accountable when something goes wrongFDA Q21
“However, distributed data collection should not result in distributed accountability.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“For GenAI, PCCPs may therefore need to rely more heavily on procedural and boundary-based prespecification rather than exhaustive technical prespecification of every future modification.”
Devices that plan and take actionsFDA Q26
“An erroneous output in one step may become the input to subsequent steps, trigger additional tool calls, or result in actions that further amplify the original error.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“Human oversight should therefore be treated as a risk control whose effectiveness may require both initial and periodic evaluation, rather than as an inherently reliable safeguard.”
Records that let investigators reconstruct an eventFDA Q19, Q21, Q24, Q26
“For meaningful investigation of complaints, adverse events, or performance signals, the manufacturer should be able to reconstruct, to the extent technically feasible, the relevant device configuration at the time of use.”
How this fits rules that already existFDA Q8, Q9, Q16, Q25
“This ambiguity becomes even clearer when considering the definition of risk in ISO 14971.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ26 · Agentic devices

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
Keep it, but add or change elements
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider the wording and specificity
Consider how personalized the answer is
Consider the user and clinical context
Q3When clinical information goes straight to the patient, does the risk change, and what safeguards help without underestimating patients?
Base risk on the task and available safeguards
Q4Should it matter whether the clinician using the AI is a generalist or a specialist?
Assess the clinician’s task-specific knowledge
Q5How is risk assessed when a conversation starts with non-directive information and drifts into action-directing?
Test whole conversations, not isolated answers
Enforce limits on what the conversation can do
Q6How should under-escalation be weighed against over-escalation?
Set the trade-off for the clinical context
Q13Where is synthetic data good enough, and where is it not?
Use synthetic cases for rare events and stress testing
Check for shared blind spots in generated test data
Keep real evidence for claims synthetic data cannot establish
Q15Could the AI be measured against what would have happened without it: unaided judgment, a delayed specialist, or no intervention?
Use that comparator, with conditions
Q18Can greater premarket uncertainty about a GenAI device’s benefit-risk profile be accepted through greater reliance on postmarket monitoring?
Allow it only under defined conditions
Q19How should an AI device be monitored after launch, and what sets the cadence?
Repeat performance testing on a schedule
Have clinicians review samples of outputs
Q20Could AI supervisory agents help carry out postmarket monitoring?
Use AI monitoring with validated safeguards
Q21What roles should clinicians and institutions play in monitoring, without diluting manufacturer accountability?
Keep the manufacturer responsible for investigation and action
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Q22With the premarket competency assessment as the baseline, which post-deployment changes need re-evaluation, and how much?
Scale retesting to the change’s clinical impact
Q23How can a change-control plan cover changes that cannot be fully specified in advance?
Define what must remain safe instead of predicting every edit
Specify the tests or controls a future change must pass
Q24When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Q26What extra oversight does an AI that plans and acts in multiple steps need?
Evaluate the full sequence of actions and its effects
Keep records that let investigators reconstruct actions
Tighten performance requirements as autonomy increases
Clarify responsibility and reportable failures

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Acts
High-consequence work: Directs
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Thank you for your efforts to clarify the regulatory background for generative AI (GenAI)-enabled medical devices and for the opportunity to provide feedback on this important topic. Please find my comments in the attached document.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

U.S. Food & Drug Administration
Center for Devices and Radiological Health (CDRH)

Dear FDA Team,

thank you for your efforts to clarify the regulatory background for generative AI (GenAI)-enabled medical devices
and for the opportunity to provide feedback on this important topic. Please find my comments and suggestions
below.

Section IV – Considerations for the Assessment of Risk for GenAI-Enabled Devices
General Considerations and Response to Discussion Question 1 – Risk Assessment
I support FDA's objective of establishing a risk-proportionate framework for the regulation of GenAI-enabled
medical devices. The proposed two-axis framework based on device activity and the consequences of relying on
an incorrect output may provide a useful and pragmatic starting point for high-level regulatory decision-making.
However, I believe that the framework would benefit from a clearer definition of the type and purpose of the
“risk assessment” being performed.
The Discussion Paper appears to focus primarily on an initial assessment used to derive regulatory requirements
for a given device. This is an important phase to consider. However, this type of risk assessment may be confused
with the product-specific risk assessment that is performed during product development and throughout the
product lifecycle. In the Discussion Paper, it seems that the term “risk” is used at both levels even though the
underlying concepts do not fully coincide. At the first level, only information that is already defined and can be
reliably supported when the initial regulatory criticality is assigned (e.g., in the device description) should be used.
Information that becomes available only through subsequent development and risk management should not yet
be relied upon. For example, reliable product-specific estimates of the probability of occurrence of harm, or of
probabilities associated with specific failure and harm scenarios, may often not yet be available at this stage.
This ambiguity becomes even clearer when considering the definition of risk in ISO 14971. In ISO 14971, risk is
defined as the “combination of the probability of occurrence of harm and the severity of that harm”. As discussed
above, product-specific probabilities may often not yet be sufficiently characterized during the initial phase.
Consequently, the term “risk” as used for the initial regulatory assessment may be interpreted differently from
the ISO 14971 definition of product risk, which could create ambiguity. In relation to the two-axis framework as
proposed in the Discussion Paper, this definition of risk overlaps with the consequences axis. Some passages of
the Discussion Paper appear to use the consequences axis more broadly by also considering factors that affect
reliance on an incorrect output or the likelihood that such an output results in harm. This seems to be more
related to product-specific risk assessment. However, this relationship is not explicitly defined.
From my perspective, the Discussion Paper should clarify more explicitly the context in which the term risk is
used. It should always be clear at which level the term operates. The two levels should be clearly separated and
the terminology should indicate the respective context. I suggest using different terms to avoid confusion: the
first level could be described as “criticality” of the product, whereas the second could remain “risk” in accordance
with ISO 14971.
Furthermore, I suggest separating the two phases and clearly delineating them. I propose the following names for
them, while recognizing that these terms are not established regulatory categories.

1. Criticality Assessment / (Regulatory) Criticality Stratification:
This phase is used to determine the appropriate level of regulatory oversight and evidence for a defined device
and use context. At this stage, only information that is already defined and can be reliably supported should be
used, e.g., the type of activity as defined in the Discussion Paper or the severity of a potential harm. Other factors,
such as the application context, the required competency profile of users (e.g., lay persons/patients versus
healthcare professionals), or other fixed parameters regarding integration of the device into the clinical workflow,
may also be considered. This phase could be called Criticality Assessment when it focuses on the general
assessment of criticality. Subsequently, I use the term Criticality Stratification because it refers to a categorization
into criticality levels that can then be used to determine the applicable regulatory requirements.

2. Product-Specific (Safety) Risk Management
This phase refers to the established risk management process according to ISO 14971 that is performed
throughout device development and the total product lifecycle, including risk analysis, risk evaluation, risk
control, and evaluation of residual risk consistent with ISO 14971.
For the following comments, I interpret Section IV of the Discussion Paper primarily as referring to the first phase,
i.e., the Criticality Stratification introduced above. The following considerations are based on this interpretation.
This initial phase sets the anchor for subsequent regulatory and development steps. Accordingly, only factors that
are defined and can be reliably supported at this stage should be used. These may include the following factors
where the first two entries primarily correspond to the two axes already suggested in the Discussion Paper. The
other entries may contribute to these axes (e.g., by affecting the consequences of outcomes) but could also be
considered as distinct modifiers or dimensions.
• Potential severity of harm that may result from relevant erroneous outputs;
• Activity and degree of autonomy, ranging from non-directive information to autonomous action-taking;
• Directness of the clinical effect, including whether an output indirectly supports clinical reasoning or directly
determines or initiates a clinical action;
• Scope restrictions, such as excluded indications, populations, actions, or clinical contexts;
• Intended user, including patient-facing versus healthcare-professional-facing applications;
• Reversibility of actions or consequences;
• Qualification, expertise, and experience of the intended user, including whether specialist-level knowledge is
required to independently evaluate an output;
• Clinical setting, such as home use, outpatient care, inpatient care, intensive care, or emergency care;
• Time criticality, including the time available for independent review, additional diagnostic investigation, or
correction before harm may occur;
• Fundamental workflow safeguards or safety boundaries that are already binding characteristics of the
defined device and intended workflow;
• Availability of independent information or review that enables an output to be assessed before reliance upon
it; and
• Traceability of an output to relevant source information.

The distinction between initial Criticality Stratification and subsequent Product-Specific Risk Management is
particularly important when considering these aspects. As an example, I use workflow safeguards to illustrate the
difference between the Criticality Stratification and Product-Specific Risk Management phases. A safeguard
should be included in the initial Criticality Stratification only where it is already an inherent or binding part of the
defined device or intended workflow. For example, a technically enforced requirement for independent clinician
confirmation before an irreversible high-consequence action may legitimately be considered as part of the initial
device and use profile. In contrast, anticipated risk controls that may be developed later should not be used to
reduce the initial regulatory criticality merely because the manufacturer expects them to be implemented during
development. This distinction would help prevent regulatory categorization from depending on hypothetical
future controls and would maintain a clear boundary between Criticality Stratification and Product-Specific Risk
Management. Where safeguards are credited at the stratification level, they should be defined, enforceable,
verifiable, clinically realistic, and maintained throughout the total product lifecycle.
Other factors, such as probability-related information or the degree of certainty/uncertainty of the available
evidence, could also be considered where such information is already sufficiently supported at this stage. This
could allow regulatory requirements to be better tailored to the technical and clinical maturity of a given device
or technology. For example, this may provide an appropriate way to account for well-established technologies for
which substantial relevant evidence is already available.
Based on these considerations, the Criticality Stratification step could be strengthened or extended. As an initial
step, it could clarify not only basic characteristics such as intended use, user profiles, and application context, but
also additional aspects of the criticality profile and the requirements or constraints needed to keep the product
within a defined criticality level. This could help address the particular complexity of GenAI-enabled medical
devices, including the additional factors identified in the Discussion Paper that may influence outcomes.
The Criticality Stratification should be based on assumptions that can be reliably supported for the defined device
and use context. Where a range of possible device behaviors or conditions exists, the stratification should
conservatively reflect the highest level of criticality that can reasonably occur and cannot be reliably excluded by
defined and enforceable constraints. Accordingly, probability-related assumptions should only be used where
they can be sufficiently supported or conservatively bounded. Otherwise, information generated later during
development may invalidate assumptions underlying the initial criticality level and indicate that the selected
regulatory pathway is no longer appropriate.
At the same time, the initial Criticality Stratification need not be considered static. If information generated
during development, validation, clinical evaluation, or postmarket use materially challenges the assumptions
underlying the initial stratification, the criticality classification and associated regulatory requirements would
need to be reassessed. This feedback mechanism would integrate the initial stratification into the total product
lifecycle while maintaining the conceptual distinction from Product-Specific Risk Management. See also the
comments in Section V regarding the interplay between Criticality Assessment and development as well as
evaluation steps.

Overall comment for Discussion Questions 2–6
Across Questions 2–6, the following principle is applied, consistent with the rationale presented above:
A factor or characteristic should influence the initial Criticality Stratification only to the extent that it is a clearly
defined and reliably assured property of the device, including intended use, intended user, or use environment.
Characteristics whose actual effectiveness can only be addressed during product development, validation, clinical
evaluation, or real-world use should instead be addressed primarily through Product-Specific Risk Management.
Most characteristics relevant to Criticality Stratification may also be relevant to Product-Specific Risk
Management. Conversely, Product-Specific Risk Management may include additional factors that cannot yet be
used for the initial Criticality Stratification, for example product-specific probabilities that are not yet sufficiently
characterized or bounded at this stage.

Response to Discussion Question 2 – Degree of Directiveness of Informational Functions
I agree that directiveness should be considered a continuum rather than a binary distinction. The following
aspects may be relevant. For each factor or characteristic, it should be considered whether it can contribute to
Criticality Stratification or can only be addressed during Product-Specific Risk Management.
• Specificity of the proposed action, e.g., general information versus a concrete instruction;
• Strength of recommendation and availability of alternatives, e.g., whether the user is presented with several
reasonable options or effectively directed toward one;
• Enforceability of safeguards – the extent to which the application of relevant safeguards can be reliably
ensured in the intended setting;
• Availability and, where applicable, requirements for independent review before action is taken;
• Time criticality/immediacy – whether the output suggests action at some future point or requires immediate
action;
• Clinical significance of the action, e.g., minor self-care versus major diagnostic or therapeutic intervention;
• Degree of personalization – generic information versus a recommendation based on the individual patient's
data.

Response to Discussion Question 3 – Patient-Facing versus HCP-Facing Functions
I agree that the distinction between patient-facing and HCP-facing functions is important and should be
considered in Criticality Stratification. However, the key consideration is the extent to which the intended use
environment allows relevant safeguards, independent review, and workflow constraints to be reliably
implemented and enforced. This distinction may be particularly important when comparing home care settings
with use in a healthcare institution.
In a healthcare institution, it may be possible to better establish and enforce:
• defined clinical protocols;
• mandatory review steps;
• access to independent clinical information;
• defined escalation procedures;
• defined user qualifications; and
• technical or organizational restrictions on subsequent actions.

In a home care environment, comparable safeguards may often be considerably more difficult to guarantee.
Instructions such as “consult your physician before taking action” may not provide the same level of assurance as
an enforced review step within a controlled clinical workflow.
Accordingly, safeguards should reduce criticality only where they are sufficiently defined, enforceable, verifiable,
and realistic for the intended environment.
The distinction should therefore be based less on the label “patient”
versus “HCP” and more on the actual degree of guaranteed independent review, user competence, workflow
control, and ability to prevent inappropriate reliance.

Response to Discussion Question 4 – Generalist versus Specialist HCPs
I agree that the distinction between generalist and specialist physicians may be relevant for Criticality
Stratification. However, several caveats should be considered.
A specialist may have substantially greater expertise regarding a specific disease, intervention, or diagnostic
question. However, a generalist may in some circumstances have a broader understanding of the patient's overall
clinical situation, including comorbidities, concurrent treatments, and effects outside one specialty domain. The
relevant consideration should therefore be whether the intended user possesses the competence and
information necessary to independently assess the particular output, rather than professional designation alone.

The intended clinical environment may also be relevant. In specialist departments, narrowly defined workflows
and condition-specific protocols can often be implemented more consistently. General practice may involve a
much broader spectrum of clinical situations, making comprehensive case-specific protocols more difficult to
establish.
Accordingly, any criticality-modifying effect should be based on the competence, information, and workflow
conditions that can actually be assured for the intended user group rather than on professional designation alone.
Response to Discussion Question 5 – Multi-Turn Conversations and Migration of Function
The Discussion Paper appropriately notes that a conversational GenAI-enabled device may migrate from nondirective information to action-directing information over the course of an interaction and proposes considering
realistic conversational trajectories rather than isolated outputs. As discussed in the answer to Question 1, a
conservative approach should be applied for Criticality Stratification. The highest criticality that can reasonably
occur within the defined use and cannot be reliably excluded by enforceable safeguards should be used as the
basic reference. The relevant criterion should therefore be the highest criticality that remains reasonably
foreseeable across such conversational trajectories.
It may be possible to use appropriate safeguards to maintain a defined level of criticality. Again, the safeguards
should be reliably enforceable and their effectiveness reasonably assured. Examples of such safeguards may
include:
• technically enforced scope restrictions;
• enforceable conversational protocols that constrain the conversation to the defined scope;
• refusal or deferral mechanisms;
• mandatory transition to an HCP;
• restrictions on certain categories of recommendations; or
• reliable human-oversight checkpoints before higher-criticality outputs can be acted upon or higher-criticality
actions can take effect.

Response to Discussion Question 6 – Under-Escalation and Over-Escalation
I agree that both under-escalation and over-escalation may result in harm and that the two types of error may not
be directly commensurable. The potential severity of harms resulting from these different error modes may
therefore influence Criticality Stratification. However, detailed weighting of the associated risks will often only be
possible at the level of Product-Specific Risk Management, where product-specific probabilities and the
effectiveness of implemented risk controls can be considered. For Criticality Stratification, only factors that can
already be reliably established or bounded at this stage should be taken into account. This may include
safeguards, such as defined escalation procedures integrated into the GenAI-enabled device, that can reliably
constrain the consequences of over-escalation or under-escalation.
I would also be cautious about treating escalation primarily as a binary problem of “escalate” versus “do not
escalate.” GenAI-enabled devices may provide a spectrum of recommendations, for example:
• continue self-management;
• arrange routine medical review;
• contact a healthcare professional within a defined period;
• seek urgent care;
• seek immediate emergency care.

Errors may therefore involve not only false-positive or false-negative escalation, but also incorrect escalation
level, inappropriate timing, or inappropriate destination of care.
The Criticality Stratification should reflect these
graded outcomes and the clinical consequences of moving a patient in either direction along this escalation
spectrum.
Section V – A Competency-Based Approach for Premarket Evaluation of GenAI-Enabled
Devices
General Considerations for Section V – including Responses to Discussion Questions 7–8
I generally support the competency-based approach proposed in the Discussion Paper, in particular the
combination of device benchmarking and clinical confirmation and the concept of tailoring the nature and rigor of
evidence to the device's intended use and criticality. Given the open-ended input and output space of GenAIenabled devices, exhaustive testing is neither feasible nor an appropriate regulatory objective.
However, I again suggest distinguishing these aspects more clearly. In particular, this applies to the already
introduced delineation between Criticality Stratification and Product-Specific Risk Management. On the one hand,
it should be demonstrated that the boundaries defined during Criticality Stratification are robustly established.
Assumptions or safeguards that materially reduce the assigned criticality or evidence burden should themselves
be supported by evidence. Thus, appropriate evaluation should be included for this. For example, where lower
criticality is assumed because a qualified healthcare professional independently reviews an output, evaluation
should establish not merely that a human approval step exists. It should validate that the intended users can
realistically identify and correct safety-relevant errors under the intended conditions of use. Findings from evaluation should be capable of feeding back into the initial criticality assessment if such assumptions prove unreliable.
Conversely, Criticality Stratification can serve as a basis for defining the required extent of premarket evaluation.
In principle, this can be aligned with the two-axis framework described in the Discussion Paper. However, I would
include additional aspects as discussed in the responses to Questions 1–6 in Section IV. Based on this, most
evaluation activities will then be directed toward product development and, in particular, Product-Specific Risk
Management. Ultimately, evaluation should provide reasonable assurance that the GenAI-enabled device is safe
and effective. In this regard, different levels of criticality could be used to shape the evaluation requirements in a
proportionate manner. This may support a staged evaluation approach in which criticality is not treated as
completely static but can be reassessed step by step as evidence develops – see further comments below.
The analogy to competency assessment of human clinicians is useful as an organizing principle, particularly with
respect to structured assessment, supervised practice, progressively greater independence, and periodic
reevaluation. The Discussion Paper itself describes the analogy as one that requires adaptation to the technical,
practical, and legal characteristics of medical devices. However, I would be cautious to overemphasize the analogy
to an assumption of equivalence between clinician and GenAI competency.
Clinical professionals operate within
an education, accountability, and professional-practice framework that differs materially from the mechanisms
underlying GenAI systems. In addition, increasing reliance on AI may alter the capabilities and behavior of the
human users themselves. For example, deskilling is an additional risk that should be considered with extensive
use of GenAI in clinical settings. Human oversight should therefore be treated as a risk control whose
effectiveness may require both initial and periodic evaluation, rather than as an inherently reliable safeguard.

Distinction between errors and consequences
Additionally, I suggest more clearly distinguishing between errors in the output of the GenAI system and the
consequences of the errors in clinical practice. In scientific publications, the evaluation of AI systems often
focuses on errors (e.g., accuracy rates) but not on the consequences of the different types of errors (e.g., false
positives or false negatives). The consequences may strongly depend on the risk management measures and
other safeguards implemented in the GenAI-enabled device or in the healthcare organization or clinical practice.
On the one hand, benchmark tests often focus more on detecting errors in the output of the GenAI device. They
may be weighted according to clinical impact where that impact can be assessed in a sufficiently generic way.
Thus, they may also include an analysis of more general consequences. On the other hand, clinical confirmation is
more directly oriented toward the outcome in a concrete clinical setting. Especially for GenAI devices whose
performance usually depends on a wide range of parameters, e.g., use environments, clinical protocols,
conversational paths, I agree that clinical confirmation is an important step that should be established separately.
Although both evaluation parameters (errors and consequences) can be addressed in both steps (device
benchmarking and clinical confirmation), I suggest clarifying these relationships. From my perspective, it is crucial
to clearly identify which parameters are best addressed in each phase.

Progressive evidence-based deployment
The concept of "supervised practice with progressively greater independence" could be extended and
operationalized more directly for GenAI-enabled devices. A device could initially operate within a narrowly
defined and highly controlled operating envelope – for example, limited indications, specialist users, mandatory
independent review, or restrictions on action-taking – and subsequently expand its scope or independence only
when prespecified evidence thresholds have been met. This would also lead to a staged approach for stepwise
expansion, e.g., through the following sequence:
→ strongly controlled conditions
(e.g., restricted to dedicated experts or strong safeguards)
→ supervised clinical use
→ expanded use
This would strengthen the analogy to HCP training, in which competence is established progressively through
different levels of supervision and responsibility.
Such an approach could use predetermined evidence gates, including prespecified performance and safety
criteria, together with criteria for restriction or rollback if performance deteriorates. Where legally and practically
appropriate, Predetermined Change Control Plan (PCCP)-like mechanisms may provide a useful model for
prespecifying some such transitions. This could create a more direct link between premarket evaluation,
supervised clinical deployment, periodic evaluation, and lifecycle change control.

Response to Discussion Questions 9–10 –
Benchmarking should evolve from datasets toward broader evaluation environments
The benchmarking elements proposed by the Discussion Paper cover many important dimensions, including
safety, clinical proficiency, generalizability, and agentic capabilities. As already mentioned, benchmark evaluation
could include metrics related to clinical consequences rather than addressing output errors alone. Usually, testing
should not only be performed against a single representative dataset. For GenAI systems, a broader evaluation
suite may be more appropriate, combining different evaluation assets for different purposes. For example, this
could include distribution-representative test sets to estimate expected performance, risk-enriched challenge sets
for rare but clinically consequential failures, prespecified subgroup or scenario-specific test sets, boundary and
out-of-scope cases, and dynamic multi-turn scenarios. Average performance alone should not permit strong
performance in common or low-risk situations to compensate for poor performance in a clinically critical subset.

For safety-critical failure modes, prespecified performance floors or separate acceptance criteria may therefore
be appropriate.
Particular attention should be given to dynamic conversational trajectories that could also be included in
corresponding benchmarks. In this regard, benchmark testing shifts from a static regime with fixed datasets
toward more dynamic testing approaches, e.g., using tools that generate simulations of conversational
trajectories. For example, simulated patients, users, or clinical environments could allow scalable assessment of
how performance changes with incomplete information, different orders of presentation, user behavior, health
literacy, clinical expertise, time pressure, or evolving clinical states. Such simulation could also test whether the
device gathers the right additional information, identifies conflicting assumptions, escalates appropriately, and
recovers from erroneous intermediate conclusions. The Discussion Paper already anticipates the possible use of
simulation tools and virtual patient avatars. This concept could be developed into a more explicit intermediate
layer between static benchmarking and clinical confirmation. However, simulation cannot fully substitute for
validation with real users where the assumed risk reduction depends on human behavior or workflow. A
simulated user may efficiently expose potential interaction failures but cannot establish that actual clinicians or
patients will behave in the same manner.
Benchmarking should also explicitly evaluate device-specific safeguards and consistency mechanisms identified
through Product-Specific Risk Management or defined within the Criticality Stratification phase. Such mechanisms
may include scope restrictions, deferral and escalation rules, human-approval checkpoints, cross-checks, or
mechanisms that prompt either the user or the device to challenge important assumptions. Where one AI
component is used to supervise another, the degree of independence and the possibility of correlated or
common-mode failure should be considered.
Finally, benchmarking may support lifecycle regression testing. In particular, this may address the topic of updates
to foundation models. Testing different suitable foundation models during development may help characterize
which device behaviors are robust to model substitution and which are highly foundation-model dependent.
Regulatory evidence should nevertheless remain focused on the deployed configuration. Where an underlying
model subsequently changes, a prespecified regression suite, particularly for safety-critical competencies and risk
controls, could provide systematic evidence of the impact of the change. Over successive updates, such testing
could also establish an empirical understanding of which device properties are sensitive to foundation-model
changes.

Response to Discussion Questions 11–13 –
Clinical confirmation, representativeness, and synthetic data
The level of clinical confirmation should be determined not only by criticality but also by the remaining
evidentiary uncertainty after benchmarking and simulation. Prospective evidence becomes particularly important
where safety depends materially on human-AI interaction, local workflow, real-time information availability, or
where use of the device changes subsequent diagnostic or therapeutic decisions.
In contrast, rare safety-critical
failure modes may sometimes be characterized more efficiently through enriched challenge testing and
simulation than through very large prospective studies.
For retrospective evaluation in particular, temporal fidelity should be included as another important topic. The
information provided to the device during testing should reconstruct the information state that would actually
have been available at the intended clinical decision point. Later or additional investigations, diagnoses,
treatment decisions, or outcome information may appropriately contribute to establishing the reference
standard, but should not inadvertently be made available to the device being evaluated. Without this separation,
retrospective testing may substantially overestimate performance. A related issue can arise during model
development when training data contain proxy variables or information that would not be available at the
intended decision point. In such cases, models may exploit shortcuts that link this additional information to the
outcome. This can compromise the intended temporal and causal structure of the clinical decision problem, since
model inputs should reflect information that would also be available for new cases.
Representativeness should also be treated as a multidimensional concept. A dataset that accurately reflects realworld prevalence may contain too few rare but high-consequence cases to establish safety with adequate
precision. Conversely, a risk-enriched challenge set is intentionally not prevalence-representative. I therefore
suggest distinguishing the evidence needed to estimate expected real-world performance from the evidence
needed to establish coverage of clinically and risk-relevant situations. Ultimately, the evaluation design should
reflect the risks of the GenAI-enabled device as it is applied in the actual clinical setting.
Statistical evaluation should correspondingly include overall performance together with prespecified subgroup,
scenario, and failure-mode-specific analyses where these are relevant to safety. Sample-size requirements should
be driven by the precision required for the relevant risk-based performance claims rather than by a uniform
minimum number of test cases. Again, the required precision should be related to the risks in the actual clinical
setting.
Synthetic data can be particularly valuable for rare conditions, counterfactual testing (what-if-scenarios),
controlled variation of patient or environmental characteristics, and generation of difficult conversational
trajectories. They should generally augment rather than replace real data, particularly where clinical data are
limited. Additionally, their regulatory role should be explicit. Synthetic data used for stress testing or coverage
expansion need not reproduce actual prevalences or risk profiles exactly. Instead, synthetic data used to estimate
real-world performance require substantially stronger evidence of such types of representativeness. Results from
these different purposes should generally not be pooled into a single overall performance estimate.
Independence of the generator, clinical plausibility, and the risk that synthetic data reproduce the same biases as
the device under evaluation should also be considered. The Discussion Paper already identified the latter concern
in Question 13.
For devices whose safety materially depends on deployment-specific conditions, e.g., user expertise, local clinical
pathways, availability of downstream safeguards, or technical integration, a proportionate form of local
deployment qualification could be appropriate. This need not imply full revalidation at every institution. But the
greater the regulatory reliance on local conditions, the stronger the evidence should be that those conditions
actually exist and function as assumed.

Response to Discussion Questions 14–15 – Reference standards and clinically relevant comparators
Regarding performance comparators, I suggest more clearly separating the role of the reference standard from
the role of the clinical comparator. Where a clinical comparator is used, it should reflect a relevant alternative to
the device rather than being conflated with the reference standard. According to standard rules for medical
devices, the GenAI device needs to be compared to the performance of the established standard-of-care to gain
market access. This means, that it needs to be assessed how the GenAI system performs in relation to the
standard-of-care when comparing both outcomes to an idealized reference standard or ground truth. The
reference standard itself should establish, as reliably and accurately as possible, what diagnosis, assessment, or
course of action is correct or clinically appropriate for the particular test case.
Instead, the comparator, i.e. standard-of-care, should represent the care that would realistically occur in the
absence of the device, including the intended user group and clinical environment. The Discussion Paper already
raises this possibility in Question 15, including unaided clinical judgment, delayed specialist review, or no
intervention. Accordingly, the experts establishing the reference standard need not be the clinicians against
whom the device is compared. For a device intended to support generalist physicians, for example, specialist
adjudication may establish the reference standard while representative generalist physicians provide the clinically
relevant comparator. Similarly, for patient-facing home-use devices, representative patients or lay users may be
the relevant user comparator, while the reference standard may still rely on specialist adjudication or another
high-quality clinical reference.
This approach shifts the focus from an absolute comparison of GenAI with an expert, who may not represent the
intended user, toward a relative assessment in which the GenAI-enabled device and the intended user group are
compared by reference to the same reference standard. Based on this, the demonstration of non-inferiority to
the comparator as a relative criterion gets the main objective instead of determining an absolute level of
deviation between the GenAI and the reference standard.
The evaluation should also generally focus on the combined human-AI system rather than comparing the standalone GenAI-enabled device with the intended user. More generally, the following clinical comparison may be
appropriate:
• intended user + standard of care + GenAI device
versus
• intended user + standard of care without the GenAI device,
with both evaluated against the same independent reference standard. This provides a more direct assessment of
whether the device improves or at least preserves decision quality within its intended context of use than a
simple comparison of stand-alone GenAI output with expert opinion. Relative non-inferiority or superiority
approaches may help avoid arbitrary absolute performance thresholds, although critical safety outcomes should
still be subject to absolute risk-based acceptance criteria.
This approach also creates a clearer distinction between the information used to establish the reference standard
and the information available to the evaluated users or device. While subsequent diagnostic findings or follow-up
may strengthen the reference standard, the investigational and comparator conditions should only have access to
information that would realistically have been available at the relevant decision point.

Overall conclusion for Section V
In summary, I support the proposed combination of competency-based benchmarking and clinical confirmation,
but suggest strengthening the framework, in particular, in the following areas.
• better discriminate between evaluating errors in the output of the GenAI device and consequences that
result from these errors;
• more explicitly link evidence intensity to criticality while using Product-Specific Risk Management to
determine the specific content of evaluation;
• expand benchmarking from predominantly dataset-based assessment toward dynamic, risk-based evaluation
suites;
• establish the use of an independent reference standard where the GenAI device can be compared to the
established standard of care in a relative way;
• distinguish intrinsic device performance from performance of the device within the intended human and
clinical system;
• treat evaluation as a lifecycle process in which critical assumptions, safeguards, and the authorized operating
envelope can be progressively confirmed, expanded, and periodically reassessed.

Such an approach would preserve the scalability sought by a competency-based framework while providing a
clearer link between device competence, risk-control effectiveness, clinical decision quality, and ultimately
reasonable assurance of safety and effectiveness.

Section VI – Postmarket Monitoring for GenAI-Enabled Devices
General Considerations for Section VI
I generally support the Discussion Paper's view that robust postmarket monitoring will be particularly important
for GenAI-enabled devices. Their open-ended outputs, potential non-determinism, dependence on evolving
foundation models and deployment environments, and capacity for post-deployment modification make it
difficult for premarket evaluation alone to provide a complete and durable characterization of safety and
effectiveness. The Discussion Paper appropriately recognizes both the limitations of premarket testing and the
potential role of periodic re-benchmarking, clinician review, and performance degradation monitoring.
From my perspective, postmarket monitoring should serve two complementary purposes. First, it should
continuously assess the product-specific safety, performance, and benefit-risk profile. Second, it should verify that
the assumptions and safeguards underlying the device's Criticality Stratification remain valid in real-world use.
This reflects the purpose of Criticality Stratification described in the responses to the previous Discussion
Questions. For example, where the acceptability of a device depends on effective human oversight, restricted
scope of use, escalation mechanisms, or other downstream safeguards, postmarket monitoring should assess
whether these safeguards remain effective in practice rather than merely whether the GenAI-enabled device
continues to meet technical performance thresholds.
Postmarket evaluation should also extend beyond isolated model performance. Static benchmarking may identify
errors in outputs but often does not capture user interaction, workflow effects, automation bias, human-AI team
performance, or downstream clinical consequences. Depending on device criticality, postmarket evaluation
should therefore progressively include dynamic interaction-based assessment, real-world use data, and clinically
meaningful outcomes. This is consistent with recognition that benchmarking alone may not fully establish
performance in real clinical use and that clinical confirmation may need to consider workflows, users,
populations, and downstream outcomes – as addressed in the Discussion Paper.
An additional challenge is that both the device and its environment may evolve. GenAI-enabled devices may
change through underlying data, model updates, guardrails, orchestration logic, or other components, while
healthcare delivery, clinical guidelines, workflows, user competencies, and patient behavior may simultaneously
change. This creates a form of co-evolution between the device and its clinical environment. Consequently,
postmarket monitoring should periodically reassess not only the device, but also whether the benchmarks,
reference standards, clinical assumptions, and deployment conditions underlying the original evaluation remain
valid. In particular, this is important due to the rapid pace of AI technology development and the corresponding
evolution of healthcare organizations.
Finally, postmarket monitoring should be linked to predefined actions. Monitoring is only effective where relevant
changes can be detected sufficiently early and where the rate of system change remains compatible with the time
required to understand and mitigate emerging risks. A monitoring strategy should therefore include defined
metrics, thresholds, triggers, escalation pathways, and, where appropriate, rollback, restriction of use, increased
human supervision, or renewed clinical confirmation. One challenge is ensuring that governance and monitoring
feedback loops operate on a timescale commensurate with system change. Otherwise, detection and mitigation
may lag behind changing conditions, and system behavior may evolve faster than it can be adequately assessed
and controlled.

Discussion Question 18 – Premarket vs. Postmarket Activities
In principle, increased reliance on postmarket monitoring can be appropriate, but it should not create a general
trade-off whereby stronger postmarket monitoring automatically permits less premarket evidence.
The
Discussion Paper specifically frames this question in terms of uncertainty regarding the device's benefit-risk
profile, which ultimately remains a product-specific determination.
The acceptable degree of premarket uncertainty should primarily depend on the device's criticality (as introduced
in the feedback to Sections IV and V) and on whether emerging risks can be detected and controlled before
clinically significant harm occurs. Relevant considerations include:
• severity and reversibility of potential harm;
• detectability and latency of emerging failures;
• degree of autonomy and effectiveness of human or technical safeguards;
• number and rate of patients exposed;
• ability to restrict use, suspend operation, or rapidly roll back changes; and
• stability of the underlying model and deployment environment.

Higher-criticality devices should generally require greater premarket assurance and more intensive postmarket
monitoring. Increased postmarket monitoring should be able to compensate for premarket uncertainty only
where the residual uncertainty is sufficiently observable, reversible, and controllable.
Postmarket monitoring should additionally verify continued compliance with the assumptions used in the
Criticality Stratification, including the actual effectiveness of human oversight, scope limitations, escalation
procedures, and other safeguards.

Discussion Question 19 – Triggers for Postmarket Reassessment
I support periodic re-benchmarking, clinician review, and degradation monitoring, but consider them
complementary rather than interchangeable approaches. Re-benchmarking alone should not be equated with
postmarket clinical performance evaluation. Static benchmarks may detect changes in model performance but
may fail to capture user-device interaction and the clinical consequences of device outputs. Depending on
criticality, postmarket monitoring should therefore also evaluate aspects such as human-AI team performance,
workflow effects, near misses, overrides, escalation behavior, and downstream clinical outcomes.
Dynamic evaluation approaches may be particularly useful for GenAI-enabled devices. These could include multiturn scenarios, evolving clinical information, simulated or real user interactions, and assessment under changing
deployment conditions. The Discussion Paper already contemplates benchmarking methodologies that can
support ongoing assessment over time and following modifications.
Reassessment should be both periodic and trigger-based. Relevant triggers may include:
• predefined time intervals;
• adverse events, complaints, near misses, or deterioration of risk- or performance-related metrics;
• changes in input distributions, interaction patterns, subgroup performance, or other logging-based signals;
• indications that assumptions or safeguards underlying the Criticality Stratification are no longer valid;
• changes in clinical workflows, user populations, guidelines, standards of care, or other deploymentenvironment parameters;
• changes to foundation models or other components; and
• major technological advances or unexpected disagreement patterns between the device, clinicians, and
supervisory systems.

Postmarket monitoring should address multiple forms of drift, including data drift, model/component drift,
concept or clinical drift, workflow/use drift, and oversight drift.
Importantly, reassessment should also determine
whether the benchmark itself remains representative and clinically valid rather than merely whether the device
continues to pass an unchanged benchmark.

Discussion Question 20 – Machine-Based Supervisory Agents
Machine-based supervisory agents could improve the scalability and responsiveness of postmarket monitoring
and may be particularly useful for continuous or dynamic evaluation. However, any supervisory agent performing
a safety-relevant monitoring function should itself be subject to fit-for-purpose validation and ongoing
surveillance. The Discussion Paper appropriately identifies evaluation and reliability of the supervisory agent as a
core consideration.
Relevant considerations for the use of machine-based supervisory agents may include sensitivity to relevant
failure modes, robustness, reproducibility, and susceptibility to specific failure modes. They could also be used to
detect relevant patterns in user behavior, identify potentially false assumptions, and analyze human-machine
interaction as well as interactions between the AI system and its environment. Particular caution may be
warranted where the device and supervisor rely on the same or closely related foundation models, training
sources, or technical architectures.
Machine-based supervision should generally form one layer of a broader monitoring system rather than becoming
a fully closed automated control loop. Periodic qualified human review should remain part of the monitoring
architecture, particularly for higher-criticality devices and for detecting novel or unanticipated failure modes that
an automated supervisor was not designed to identify.

Discussion Question 21 – Roles and Stakeholders
I agree that clinicians, healthcare institutions, professional societies, standards organizations, and other
stakeholders may provide important information for postmarket monitoring, particularly regarding real-world
workflows, emerging clinical practice, user behavior, and clinical outcomes. However, distributed data collection
should not result in distributed accountability.
The manufacturer should remain responsible for defining the
monitoring strategy, integrating relevant information from the ecosystem, and implementing necessary corrective
or preventive actions.
Healthcare institutions may have an especially important role in identifying changes in the deployment
environment, including workflow modifications, local protocols, user qualifications, and changes in the extent or
effectiveness of human oversight. The associated interfaces and responsibilities should be clearly defined,
including through appropriate agreements where applicable. Professional societies may contribute to identifying
changes in clinical standards, guidelines, and appropriate reference standards. For higher-criticality applications,
structured mechanisms for information exchange between manufacturers and deploying institutions may
therefore be appropriate.

Discussion Question 22 – Extent of Postmarket Monitoring Activities
In principle, the premarket competency-based assessment can serve as a baseline for evaluating post-deployment
modifications, as proposed by the Discussion Paper. However, the scope of reassessment should depend not
merely on the technical magnitude of a change, but on its potential impact on competencies, criticality,
safeguards, clinical performance, and the validity of the original evaluation assumptions. A seemingly small
technical modification may warrant substantial reassessment if it affects a safety-critical competency, humanoversight mechanism, boundary behavior, or clinical decision pathway. Conversely, some changes may reasonably
remain within the quality management system if their potential impact is well characterized and bounded.
Reassessment should also consider cumulative change. Multiple individually minor modifications may collectively
produce a substantially different device or deployment configuration.
Where relevant, post-change evaluation should preserve temporal fidelity: clinical cases should be evaluated
using only information that would have been available at the relevant decision point, while later information may
appropriately contribute to establishing the reference standard. In particular, this is important when
retrospectively assessing data from clinical cases.

Discussion Question 23 – PCCP Concepts and their Use for Postmarket Monitoring
PCCPs are an important approach for AI-enabled devices and it is reasonable to extend this concept to GenAIenabled systems. However, it should be recognized that the exact nature of all future modifications may not be
fully predictable. For GenAI, PCCPs may therefore need to rely more heavily on procedural and boundary-based
prespecification rather than exhaustive technical prespecification of every future modification.
For example, this
applies to the assumptions about application environment, user interaction, or safeguards, as defined during
Criticality Assessment or addressed during Product-Specific Risk Management.
The competency-based approach could provide a useful but extended structure for such PCCPs. In comparison to
currently pursued PCCP approaches, the stages from restricted or closely supervised use toward greater
operational independence may be integrated into this approach. Progression between these stages should occur
only when predefined competency and real-world performance criteria have been met. Conversely, emerging
postmarket signals may lead to a downgrading, such as reducing scope or autonomy or increasing required
human supervision.
PCCPs should also account for changes in the environment and the continued validity of the evaluation
framework. A device may remain technically unchanged while new clinical guidelines, workflows, user
competencies, or other environmental factors invalidate assumptions underlying its prior assessment.

Discussion Question 24 – Changes of Third-Party Foundation Models
Third-party foundation models present a particularly significant challenge because safety-relevant changes may
occur outside the direct control of the device manufacturer. An assessment of changes in the underlying
foundation models should be included explicitly in the postmarket monitoring strategy. This may include specific
benchmarking tests to assess the impact of changes in the foundation model on resulting clinical performance.
Additionally, these steps should address effects on the interaction between the GenAI-enabled device and users
as well as the deployment environment. This should be aligned with the competency-based approach already
discussed. Major changes in foundation models or other components may necessitate a downgrading of the
currently assumed competency level. The GenAI-enabled device may need to return to a stage with greater
supervision. This would amount to a requalification period in a more controlled environment. A more
independent use of the GenAI system can be re-established when defined evidence gates are successfully passed.
An additional requirement should be case-level reconstructability. For meaningful investigation of complaints,
adverse events, or performance signals, the manufacturer should be able to reconstruct, to the extent technically
feasible, the relevant device configuration at the time of use.
This may include the foundation-model version,
prompts, guardrails, retrieval sources or knowledge-base version, tools, orchestration logic, relevant runtime
parameters, and applicable clinical-environment information. Exact reproduction of every stochastic output may
not always be feasible. Nevertheless, the historical system configuration should remain sufficiently traceable to
permit reliable investigation and reassessment of safety-critical behavior.

Concluding remarks to Section VI
Overall, postmarket monitoring for GenAI-enabled devices should be understood as an active lifecycle control
system rather than a periodic performance check. Its purpose should be to maintain assurance of the productspecific benefit-risk profile, verify continued effectiveness of the safeguards underlying the device's criticality,
detect changes in both the device and its clinical environment, and trigger proportionate corrective action.
Importantly, the rate of permitted technological and operational change should remain compatible with the
ability of the monitoring and regulatory system to detect, understand, and control its consequences.

Section VII – Other Topics
Discussion Question 26 – Agentic AI
Agentic GenAI-enabled devices introduce additional system-level risks that go beyond those associated with
individual GenAI-enabled functions. In particular, their ability to autonomously plan and execute multi-step tasks,
interact with external tools, and take actions across multiple system components can make both the attribution
and control of safety-critical failures substantially more difficult. This is especially relevant when different
components are developed, maintained, or controlled by different manufacturers or organizations. In such cases,
technical and operational responsibilities as well as accountability may be distributed across different
components and providers. Risk contributions may arise across components and interfaces and therefore need to
be assessed and managed at the level of the overall system, including through Product-Specific Risk Management
and Criticality Stratification.
Existing medical device regulatory frameworks rely on an identifiable manufacturer or sponsor being responsible
for the safety and effectiveness of the device. In agentic architectures, however, the technical cause of a failure,
operational control over the relevant component, and regulatory responsibility may become separated across the
device manufacturers, foundation model providers, agent or orchestration providers, and healthcare institutions.
Nevertheless, accountability should not simply be diffused across the ecosystem. Rather, there should be a clearly
defined responsibility for the safety and effectiveness of the end-to-end system, supported by appropriate
contractual, technical, and quality-management arrangements among the parties involved. This would be
consistent with broader consideration of stakeholder involvement without diffusing manufacturer accountability,
as addressed in the Discussion Paper.
A second consideration is failure propagation within multi-step workflows. An erroneous output in one step may
become the input to subsequent steps, trigger additional tool calls, or result in actions that further amplify the
original error.
Premarket evaluation should therefore not only assess individual agents, models, or tools, but also
representative end-to-end workflows, including safety-critical interaction pathways, foreseeable failure
combinations, recovery behavior, and escalation to human oversight. This is particularly important for irreversible
or high-consequence actions.
Accordingly, testing should address system-level behavior rather than primarily focusing on individual
components. Such testing will, however, be particularly challenging because agentic systems may dynamically
select tools, alter task sequences, and operate across a large number of possible system states. Comprehensive
testing of all combinations is unlikely to be feasible. Evaluation should therefore focus on risk-based coverage of
safety-relevant configurations and representative interaction trajectories, complemented by postmarket
monitoring capable of identifying previously unobserved failure patterns. The emphasis on the final user-facing
device in its intended deployment configuration, rather than isolated subcomponents, is particularly important
for agentic systems. Responsibilities for integration testing should be clearly allocated. Component providers may
supply evidence for their components, but the party responsible for the final user-facing device should ensure
that end-to-end evidence is sufficient. Deploying healthcare organizations may need to address site-specific
integration and workflow conditions. Where applicable, a dedicated system integrator may take responsibilities
for key aspects of the integration tests.
In addition, agentic systems require interoperability at several levels. Technical interoperability is necessary to
ensure reliable exchange of data and commands, but is not sufficient. Semantic interoperability is also critical:
participating agents and tools need to interpret clinical concepts, units, timestamps, uncertainty, patient context,
and workflow states consistently. Finally, appropriate governance and regulatory alignment are required to define
responsibilities for risk management, change control, validation, monitoring, and corrective actions across
organizational boundaries. These interoperability challenges already exist in conventional interconnected medical
device environments but are likely to become substantially more complex where system composition and
behavior can change dynamically.
Finally, the reduced opportunity for direct human review should be reflected in progressively stricter acceptance
criteria as autonomy and potential severity of harm increase. Such criteria could include predefined humanoversight checkpoints for high-consequence actions, reliable detection and handling of erroneous or unavailable
tools, traceability of relevant decisions and actions, effective containment of failures, and demonstrated ability to
revert to a safe state or defer to human control. Postmarket oversight should similarly address not only
performance of individual components, but also changes in interaction behavior across the overall agentic
system.
Overall, agentic AI should not merely be treated as a more capable form of GenAI. Its autonomous orchestration
of multiple decisions, components, conversational pathways, and actions creates additional system-level risks
that require end-to-end accountability, pathway-based evaluation, effective interoperability, and riskproportionate controls on autonomy.

Concluding remarks
Taken together, these comments support a risk-proportionate and lifecycle-based regulatory approach that
clearly distinguishes the initial regulatory stratification of a device from Product-Specific Risk Management.
Criticality should help determine the level of oversight and evidence, while Product-Specific Risk Management
and evaluation should establish whether the final user-facing device, in its intended use environment, provides
reasonable assurance of safety and effectiveness.
At the same time, the framework should remain sufficiently adaptive to address open-ended behavior, evolving
foundation models, human-AI interaction, and agentic functions without treating technological novelty itself as a
proxy for risk. The key regulatory objective should be to define and verify the conditions under which a GenAIenabled device can operate safely, to detect when those conditions no longer hold, and to ensure clear
accountability and proportionate corrective action across the total product lifecycle. Such an approach could
preserve FDA's least-burdensome objective by concentrating evidence and oversight where they are most
relevant to safety and effectiveness.
Thank you again for the opportunity to provide feedback on this important topic. Please feel free to contact me if
you have any questions or would like to discuss any of these comments further.
Martin Haimerl
email: Martin.Haimerl@hs-furtwangen.de
Furtwangen University
Scientific Director Innovation and Research Center of Furtwangen University

Disclosure regarding the use of generative AI
Generative AI tools were used as assistive tools in developing, structuring, and editing this submission. The final
content was critically reviewed, revised, and approved by the author, who assumes full responsibility for all
statements and recommendations contained herein.