S. Joseph Sirintrapun, MD, FCAP, FASCP (Mass General Brigham, personal capacity)
“Scale evidence requirements to clinical consequence, autonomy, and the opportunity for meaningful human review.”
What they argued
M1 from his section 10 on agentic AI: the framework should separate an agent's ability to plan an action from its authority to execute it, and higher-consequence activities need explicit human authorization checkpoints before irreversible actions unless the device has been specifically evaluated and authorized for autonomous execution; he does not discuss patient-facing devices specifically. M2 from his concluding principle 3, stated without qualification. M4 from sections 2-4: supports competency-based evaluation of the deployed system and benchmarking followed by clinical confirmation, but conditions it - benchmarks must not become proxies for clinical validity, confirmation rigor must be proportional to risk, and human-AI team performance should be permitted as the comparator. M5 from sections 6-7: supports risk-based reassessment and PCCPs that describe categories and acceptance criteria rather than every future modification, and voluntary Foundation Model MAFs provided they do not constitute authorization of the model or shift responsibility from the device manufacturer. autonomy_high is direct because human authorization is his default before irreversible or clinically consequential actions; he leaves act-level open only where autonomous execution has itself been demonstrated and authorized, and states no level for low-consequence functions. No FDA question numbers are cited anywhere; he addresses sections instead. Type note: he is a practicing pathologist and informatician writing in a personal capacity, with Mass General Brigham and Harvard Medical School affiliations given for identification only, so clinician rather than academic or healthsystem is arguable.
Themes it raises
Across the five cross-cutting questions
High-consequence work: Directs
The comment as filed
See attached file(s)
Attachment
S. Joseph Sirintrapun, M.D. Telephone: 314-435-2432
Clinical Director of Digital Pathology at Mass General Brigham E-mail: ssirintrapun@mgh.harvard.edu
September 8, 2026
Dockets Management Staff (HFA-305)
Food and Drug Administration
5630 Fishers Lane, Room 1061
Rockville, MD 20852
Re: Docket No. FDA-2026-N-7874 - Considerations for the Regulation of Generative AI-Enabled Medical
Devices
Comments on FDA Discussion Paper: Considerations for the Regulation
of Generative AI-Enabled Medical Devices
Thank you for the opportunity to comment on the FDA Center for Devices and Radiological Health discussion
paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices. I submit these comments
in my personal capacity and reflect on my experience as a pathologist and pathology informatician involved in
digital pathology, clinical informatics, artificial intelligence, and implementing digital diagnostic systems within
large healthcare organizations.
I strongly support FDA's effort to develop a risk-based, total product life cycle approach to generative artificial
intelligence (GenAI)-enabled medical devices. The discussion paper appropriately recognizes that GenAI
presents characteristics that challenge conventional device evaluation models, including variable outputs,
open-ended inputs, multimodal data, third-party foundation models, evolving system components, and
increasingly agentic behavior.
At the same time, I encourage FDA to preserve an important principle articulated within the discussion paper
itself: regulation should follow the intended clinical function, patient risk, and degree of autonomy - not the
technology label 'generative AI.' This principle will become increasingly important as the distinctions among
predictive AI, generative AI, multimodal models, foundation models, and agentic systems become less clear.
Regulatory frameworks should remain sufficiently technology-neutral to accommodate continued innovation
without requiring fundamental reconsideration whenever model architectures change.
1. Risk should reflect clinical function, consequence, autonomy, and the opportunity for
meaningful human review
FDA's proposed two-axis framework - device activity and the consequences of relying upon an incorrect output
- is a useful starting point. The progression from informational or non-directive functions through actiondirecting, supervised action-taking, and fully autonomous action-taking captures an important dimension of
risk.
However, two additional considerations should be explicit. First, treat meaningful human review as a major
modifier of risk. The presence of a healthcare professional should not by itself qualify as risk mitigation. The
relevant question is whether the professional can realistically recognize an erroneous output before patient
harm occurs.
This distinction matters especially in pathology and other diagnostic specialties. An AI system that highlights
potentially abnormal regions of a whole-slide image for pathologist review operates differently from a system
that independently assigns a diagnosis. Yet nominally placing a pathologist 'in the loop' should not
automatically lower risk if the workflow encourages automation bias or makes independent evaluation
impractical.
Second, traceability to source information should be considered explicitly. A GenAI system that summarizes or
synthesizes information while allowing a clinician to inspect the underlying laboratory result, image, citation, or
other primary data may present a substantially different risk from an otherwise identical system whose output
cannot be independently verified.
Thus, risk assessment could reasonably consider at least four interacting factors: clinical consequence,
degree of autonomous action, opportunity for meaningful independent review, and traceability to source
evidence. These factors need not become four formal regulatory axes, but they should inform classification
and evidentiary expectations.
2. The unit of evaluation should be the deployed clinical system, not simply the foundation
model
I strongly agree with FDA's proposal that competency-based evaluation should assess the final user-facing
device as configured for real-world deployment, rather than the foundation model standing alone.
Clinical performance results from an entire system: the underlying model, prompts, retrieval mechanisms,
guardrails, software interfaces, clinical data inputs, workflow integration, and human interaction. The same
foundation model can therefore produce substantially different clinical risk profiles depending upon how it is
implemented.
This distinction is analogous to an important principle in clinical laboratory practice: vendor validation and local
clinical validation serve complementary but different purposes. A manufacturer establishes the general safety
and performance characteristics of its product. The deploying healthcare organization must determine that the
system performs appropriately within its local environment, including its interfaces, hardware, workflow, patient
population, users, and intended clinical application.
FDA should preserve this distinction rather than inadvertently creating a regulatory model in which successful
manufacturer testing is interpreted as establishing universal performance across all deployment environments.
3. Competency-based assessment is promising, but benchmarks should not become proxies
for clinical validity
FDA's proposed combination of device benchmarking followed by clinical confirmation is conceptually strong.
The benchmarking framework - including safety, clinical proficiency, generalizability, and additional
competencies for agentic systems - captures many of the characteristics that conventional fixed-output
performance testing may miss.
However, benchmarks should be regarded as tools for characterizing competency rather than substitutes for
demonstrating clinical performance. Public benchmarks may suffer from contamination, saturation, limited
representation of real-world populations, and optimization to the benchmark itself. Sponsor-developed
benchmarks introduce a different concern: a developer may knowingly or unknowingly optimize a system
toward its own test environment.
Clinical confirmation therefore remains essential, with its rigor proportional to risk. FDA's proposed spectrum -
from retrospective evaluation and shadow deployment through clinician adjudication and prospective clinical
trials - is particularly useful. A rigid requirement for prospective trials would unnecessarily impede lower-risk
innovation, while permitting benchmarking alone for high-risk autonomous functions could provide insufficient
assurance of real-world safety.
4. Evaluation should increasingly consider the human-AI team
For assistive systems, the clinically relevant unit of performance is often neither the AI alone nor the clinician
alone, but the human-AI team. This distinction is especially important in diagnostic medicine.
A system may perform below an expert physician when evaluated independently yet meaningfully improve the
performance, consistency, or efficiency of physicians using it. Conversely, a system with excellent standalone
performance may degrade clinical performance if it introduces automation bias, workflow distraction,
inappropriate reliance, or inefficient false-positive review.
Where the intended use involves professional oversight, FDA should therefore permit - and in appropriate
circumstances encourage - evaluation of clinician alone versus clinician plus AI, rather than requiring the AI
system to outperform the clinician independently. Clinical responsibility should remain with the qualified
healthcare professional when the intended use specifies an assistive rather than autonomous function.
5. Validation must become lifecycle governance
One of the strongest aspects of the discussion paper is its recognition that GenAI evaluation cannot end at
market authorization. Validation should not be treated as a one-time event. GenAI-enabled devices require
lifecycle governance, including initial validation, post-deployment monitoring, documented change control,
periodic reassessment, and defined responses to meaningful performance changes.
I therefore support FDA's consideration of periodic benchmarking, sample-based clinician review, and
monitoring for performance degradation. Postmarket monitoring should assess not only traditional accuracy
but also clinically meaningful changes in performance and calibration; subgroup performance; failure modes
and unexpected behavior; input or population drift; workflow and user interaction; automation bias; model or
foundation-model changes; and safety-critical behaviors such as escalation, deferral, and refusal.
Monitoring requirements should nevertheless remain proportional to clinical risk. Continuous monitoring does
not necessarily mean continuous regulatory submission.
6. Software changes should trigger risk-based reassessment, not automatic complete
revalidation
FDA appropriately recognizes that GenAI-enabled systems may change through software updates, retraining,
changes in prompts or orchestration, modifications to foundation models, or adaptive behavior after
deployment. The regulatory response to these changes should be proportional to their potential clinical impact.
Changes that do not materially affect intended use, clinical performance, interpretive output, data handling,
interfaces, or clinical workflow should generally be manageable within an appropriate quality management
system and should not automatically require complete revalidation.
Changes that could affect analytical or interpretive performance, patient-specific outputs, safety-critical
behavior, interoperability, or clinical decision-making should undergo documented risk-based reassessment
before clinical implementation. Depending upon the nature of the change, this might range from limited
verification or re-benchmarking to more extensive clinical revalidation.
Predetermined Change Control Plans are therefore particularly important for AI-enabled devices, but FDA
should permit them to describe categories, boundaries, evaluation processes, and acceptance criteria for
anticipated change, rather than requiring developers to predict every future technical modification precisely.
7. Third-party foundation models require transparency and contractual accountability
FDA correctly identifies an increasingly important problem: the manufacturer of a medical device may not
control the foundation model the device depends on. I support exploration of voluntary Foundation Model
Device Master Files. They could reduce redundant evaluation and provide FDA with information that individual
device manufacturers may not otherwise possess.
At minimum, foundation-model developers supporting regulated medical applications should provide
downstream manufacturers with sufficient information regarding model versioning, material updates, known
limitations, clinically relevant performance characteristics, safety mechanisms, data provenance where
appropriate, cybersecurity considerations, and advance notification of changes that could affect device
performance.
However, a Foundation Model Device Master File should not constitute FDA authorization of the foundation
model itself, nor should it transfer responsibility away from the manufacturer of the final medical device.
8. Postmarket oversight should recognize shared ecosystem participation without diffusing
accountability
FDA appropriately recognizes that manufacturers, healthcare institutions, clinicians, professional societies,
standards organizations, and other stakeholders will all participate in monitoring GenAI-enabled devices.
Healthcare organizations are particularly well positioned to detect performance changes that may emerge only
after deployment across different patient populations, workflows, and information systems.
This creates an opportunity for a distributed postmarket learning system in which healthcare organizations and
professional societies contribute structured real-world performance information. However, shared monitoring
must not become shared ambiguity regarding responsibility.
Manufacturers should remain accountable for device performance. Healthcare institutions should remain
responsible for appropriate local implementation, validation, governance, user training, and monitoring.
Qualified clinicians should remain responsible for clinical decisions where the device is authorized for assistive
use. Clearly delineating these responsibilities will be essential.
9. Interoperability, provenance, cybersecurity, and auditability should be explicit lifecycle
requirements
The discussion paper appropriately focuses on clinical performance, but GenAI systems increasingly operate
within interconnected healthcare information environments. Safety therefore depends not only on model
performance but also on the integrity of the information the system receives and generates.
FDA should encourage standards-based interoperability and lifecycle controls that preserve confidentiality,
integrity, availability, provenance, and auditability of clinical information. For GenAI and especially agentic
systems, auditability should include sufficient documentation of relevant inputs, outputs, model and version
information, tool interactions, and consequential actions to permit reconstruction of clinically significant events.
This becomes particularly important when an AI system retrieves information from external sources or acts
across multiple clinical information systems.
10. Agentic AI requires explicit control over irreversible actions
FDA appropriately recognizes that agentic systems introduce risks beyond those associated with systems that
merely generate information. The regulatory framework should distinguish between an agent's ability to plan
an action and its authority to execute that action.
For higher-consequence activities, systems should include explicit human authorization checkpoints before
irreversible or clinically consequential actions unless the device has been specifically evaluated and
authorized for autonomous execution.
Evaluation should also examine the complete chain of action rather than only the correctness of the agent's
final output. This should include tool selection, intermediate reasoning or state where technically available and
appropriate, response to erroneous tool outputs, recovery from system failures, permissions, escalation, and
the ability to terminate an unsafe sequence.
Conclusion
FDA's discussion paper represents an important evolution from regulating relatively static software toward
overseeing dynamic clinical systems throughout their lifecycle. Its proposed combination of risk-proportionate
oversight, competency-based evaluation, clinical confirmation, postmarket monitoring, and structured change
control provides a promising foundation.
As this framework develops, I encourage FDA to preserve the following overarching principles:
1. Regulate clinical function and risk rather than AI technology labels.
2. Evaluate the deployed system rather than the foundation model in isolation.
3. Scale evidence requirements to clinical consequence, autonomy, and the opportunity for meaningful human
review.
4. Evaluate human-AI team performance when that reflects the intended clinical use.
5. Treat validation as lifecycle governance rather than a one-time event.
6. Apply risk-based reassessment to software and model changes rather than automatic complete revalidation.
7. Maintain clear manufacturer accountability while enabling healthcare institutions and professional societies
to participate in postmarket learning.
8. Require sufficient interoperability, provenance, cybersecurity, and auditability to understand and reconstruct
clinically consequential AI behavior.
9. Preserve meaningful human oversight for consequential actions unless autonomous performance has itself
been appropriately demonstrated.
10. Keep regulation technology-neutral and sufficiently adaptable to accommodate model architectures that do
not yet exist.
The objective should not be to make GenAI fit regulatory constructs designed for static software, nor to create
an entirely separate regulatory system simply because a device uses generative AI. The better approach is to
evolve existing risk-based device regulation toward a lifecycle model centered on clinical function, measurable
performance, accountable deployment, and patient outcomes.
Such an approach can provide meaningful safeguards while allowing beneficial technologies to evolve without
requiring the regulatory framework to be rewritten with each new generation of AI.
Respectfully submitted,
S. Joseph Sirintrapun, MD, FCAP, FASCP
Clinical Director of Digital Pathology
Mass General Brigham
Associate Professor of Pathology
Harvard Medical School
Submitted in a personal capacity; institutional affiliations are provided for identification only.