Walnut Hill Medical
“Payers will not reimburse what they cannot define.”
What they argued
Q11 RCTs only for autonomous Class III; Q18 endorsed if monitoring plan prespecified; Q26 human confirmation before irreversible actions, action budgets.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
Coded positions
Base risk on the task and available safeguards
Test with the intended clinician group
Enforce limits on what the conversation can do
Require prospective studies for specified higher-risk uses
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Reassess after changes or safety signals
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Limit or test what the agent is allowed to do
Across the five cross-cutting questions
High-consequence work: Directs
The comment as filed
Walnut Hill Medical (WHM) is a healthcare reimbursement and commercialization strategy consulting firm with more than twenty years of engagement across medical device categories including cardiac, neuromodulation, vascular, implantable therapies, wearables, diagnostics, and surgical devices. We represent medical device manufacturers at every stage of commercialization.
We submit this comment to provide the perspective of medtech industry stakeholders on FDA’s proposed regulatory framework for generative AI-enabled medical devices. Our full comment is attached and addresses all 26 discussion questions posed in the Discussion Paper. Key positions include: (1) support for the competency-based evaluation framework; (2) the need for a published regulatory pathway decision matrix prior to framework finalization; (3) urgent interagency coordination between FDA and CMS on aligned evidentiary standards for coverage and reimbursement; and (4) practical modifications to the risk framework, benchmarking standards, PCCP change categories, and postmarket monitoring expectations. We respectfully urge FDA to prioritize regulatory clarity and least-burdensome principles as this framework develops.
Attachment
FORMAL PUBLIC COMMENT — FDA DOCKET NO. FDA-2026-N-7874
Submitted via: regulations.gov Date: August 18, 2026
To: FDA Dockets Management Staff Office of the Commissioner U.S. Food and Drug
Administration 5630 Fishers Lane, Room 1061 Rockville, MD 20852
Re: Docket No. FDA-2026-N-7874 — "Considerations for the Regulation of Generative AIEnabled Medical Devices: Discussion Paper and Request for Feedback"
From: Walnut Hill Medical (WHM) Healthcare Reimbursement & Commercialization Strategy
Consulting Dallas, TX channa@walnuthillmedical.com
I. INTRODUCTION
Walnut Hill Medical (WHM) is a healthcare reimbursement and commercialization strategy
consulting firm with more than twenty years of direct engagement across the full spectrum of
medical device categories — including cardiac devices, neuromodulation systems, vascular
therapies, implantable technologies, wearables, diagnostics, and surgical instruments. Our client
base spans early- stage medical device startups through established medtech companies, and
our work sits at the intersection of regulatory strategy, clinical reimbursement policy, payer
relations, coding and coverage analysis, and commercial implementation.
We submit this comment not as a regulatory law firm or academic institution, but as
practitioners who work daily alongside the manufacturers, clinical teams, hospital
administrators, and payer organizations that will be directly affected by FDA's regulatory
framework for generative AI-enabled medical devices. We have trained hundreds of medical
device representatives, clinical educators, prior authorization staff, billers and coders, and
hospital and ambulatory surgery center teams. We understand what happens when regulatory
frameworks create implementation friction — and when they enable innovation.
WHM's overall position is strongly supportive of FDA's collaborative approach in publishing this
Discussion Paper and requesting structured stakeholder input before finalizing regulatory policy.
The competency-based framework described in Sections V and VI represents a conceptually
sound and genuinely innovative approach. We urge FDA to continue in this spirit of
transparency, recognizing that generative AI-enabled medical devices represent a genuine
technological inflection point that demands new regulatory thinking rather than retrofitted
legacy frameworks.
Our central recommendation, developed throughout this comment, is that FDA must pair the
competency-based premarket framework with three structural commitments: (1) publication of
a clear pathway decision matrix before the framework is finalized, so manufacturers can make
rational development investments; (2) formal interagency coordination with CMS to align
clinical evidentiary standards with coverage and coding policy; and (3) a dedicated presubmission consultation mechanism for generative AI device sponsors, with particular
accommodation for small and emerging manufacturers.
II. GENERAL OBSERVATIONS
A. Competency-Based Evaluation: Conceptually Right
The Discussion Paper's central innovation — evaluating generative AI devices against defined
competencies rather than fixed performance thresholds — is analytically sound and reflects how
clinical expertise itself is validated. Just as a board-certified cardiologist is assessed not by a
single test score but by demonstrated competency across a spectrum of clinical scenarios, a
generative AI-enabled diagnostic tool should be evaluated against structured performance
domains that reflect the full range of clinical contexts in which it will operate. WHM strongly
endorses this organizing principle and encourages FDA to develop it with precision and rigor.
B. The Critical Gap: Reimbursement and Coverage Alignment
The most significant structural omission in the Discussion Paper is the absence of any discussion
of reimbursement and coverage alignment. FDA clearance or approval does not create
reimbursement rights. CMS currently operates without a systematic approach to coding and
coverage for generative AI-enabled medical devices, and commercial payers follow CMS signals
closely. If FDA clears a generative AI device without parallel work at CMS to establish a coding
pathway, coverage criteria, and evidentiary standards for coverage with evidence development,
the cleared device will face a market access valley of death that no manufacturer — and no
patient — benefits from.
Payers will not reimburse what they cannot define. They will not cover devices for which no
billing code exists, and they will apply stringent "experimental and investigational" designations
to generative AI tools until clinical evidence standards are harmonized between FDA and CMS.
The competency-based evidence FDA requires for clearance should map directly — and by
design — to the clinical evidence CMS and commercial payers require for coverage decisions.
This alignment does not happen automatically. It requires formal interagency coordination, and
FDA should commit to establishing it as part of this regulatory initiative.
C. International Harmonization Gaps
FDA's framework will not exist in isolation. The European Union's AI Act and the Medical Device
Regulation together create a parallel, in some respects more prescriptive, regulatory
environment for AI-enabled medical devices. Health Canada, Australia's TGA, and IMDRF are
each developing or refining comparable frameworks. Without deliberate harmonization — at
minimum, mutual recognition of benchmarking evidence — global medtech manufacturers will
face serial validation exercises across jurisdictions, multiplying the development burden and
creating significant competitive disadvantage relative to manufacturers operating in less
rigorous markets. FDA should treat international harmonization as a first-order priority, not an
afterthought.
D. A Tiered Pathway Map Is Essential
The Discussion Paper describes regulatory concepts without providing a clear map of which
generative AI device type belongs in which regulatory pathway. Before manufacturers can make
rational development investments, FDA must publish a decision matrix that links device risk tier
(derived from the two-axis framework) to the applicable regulatory pathway — 510(k), De Novo,
or PMA — and to the corresponding premarket evidence requirements. Absent this clarity,
responsible manufacturers will over-invest in evidence generation out of regulatory uncertainty,
and less responsible actors will under-invest and submit inadequate applications, taxing FDA
review resources.
III. SECTION IV: RISK ASSESSMENT — RESPONSES TO DISCUSSION QUESTIONS 1–6
The two-axis framework — characterizing device activity on a spectrum from non- directive to
action-directing to action-taking, crossed with an assessment of the consequences of incorrect
output — is the right organizing principle for generative AI risk stratification. WHM supports it
strongly and offers the following refinements.
Response to Question 1: Additional Dimensions of Risk
The two-axis framework captures the core of the risk landscape but requires three additional
dimensions to be complete.
First, third-party model dependency is a material risk factor not reflected in the current
framework. A generative AI device built entirely on a proprietary, auditable, manufacturercontrolled model presents a fundamentally different risk profile than a device built on an
opaque third-party foundation model over which the manufacturer has no visibility, no
contractual governance rights, and no notification of changes. The current framework treats
these devices identically. They are not identical. FDA should add third-party model dependency
as an explicit risk dimension, with devices heavily dependent on non-disclosed, externally
controlled foundation models receiving higher baseline risk classification.
Second, reversibility must be treated as a first-order risk factor. An AI- generated
recommendation that a clinician reviews and acts upon over the course of hours is materially
different from an AI-generated command that triggers an immediate, irreversible action — a
drug infusion, a neurostimulation parameter change, a surgical robot maneuver. The current
framework's consequence axis captures severity but does not fully capture irreversibility. A
separate reversibility dimension would strengthen the framework's clinical fidelity.
Third, temporal urgency — whether the device operates in real-time versus asynchronous
clinical workflows — creates meaningfully different risk profiles that the current framework
does not distinguish. Real-time autonomous action under time pressure (e.g., AI-guided
emergency drug dosing) warrants higher scrutiny than asynchronous AI-generated care plan
recommendations reviewed by a clinician before acting.
Response to Question 2: Directiveness as a Spectrum — The Need for Safe Harbor
WHM supports characterizing directiveness as a spectrum rather than a binary, but
manufacturers require a safe harbor bright line to plan development programs. Without at least
one clear presumptive standard, manufacturers face perpetual uncertainty about where their
device falls on the continuum. We recommend that FDA establish a rebuttable presumption: any
generative AI output that includes specific numeric parameters — drug doses, device therapy
thresholds, surgical dimensions — is presumptively action-directing, regardless of framing
language. Manufacturers who believe their device should be characterized differently can rebut
the presumption through pre-submission consultation. This approach provides clarity while
preserving flexibility.
Response to Question 3: Patient-Facing vs. HCP-Facing Applications
FDA is right to treat patient-facing and HCP-facing applications differently. Patients generally
lack the clinical training to contextualize AI-generated outputs, assess limitations, or override
recommendations that conflict with their broader clinical picture. WHM supports differential
regulatory treatment.
However, FDA must be careful not to overreach. Patient-facing informational tools — symptom
checkers, post-discharge medication reminders, general wellness education — should not be
regulated as action-directing devices merely because a patient reads and acts upon them. The
current framing of "action-taking" risk in patient-facing contexts risks creating a category of
regulation so broad that it encompasses general health information technology entirely. FDA
should define clear boundaries: patient-facing tools that include specific clinical
recommendations (e.g., "your symptoms indicate you should go to the emergency room") are
meaningfully different from those that provide general information without clinical specificity
(e.g., "these are common side effects of your medication"). The former warrants closer scrutiny;
the latter generally should not.
Response to Question 4: Generalist vs. Specialist HCP-Facing Applications
The distinction between generalist and specialist HCP-facing applications is clinically significant.
A generative AI tool designed for use by a board- certified interventional cardiologist carries
different risk implications than the same tool used by a primary care physician without specialty
training in interpreting its outputs. FDA should address this distinction explicitly in its labeling
and intended use requirements, requiring manufacturers to specify the intended clinical user's
competency level and to validate performance within that specific user population.
Response to Question 5: Multi-Turn Conversations — Assess at the Interaction Level
Risk in multi-turn conversational AI cannot be assessed at the level of individual turns. A single
exchange that begins with a non-directive inquiry ("What are the risk factors for atrial
fibrillation?") may, through a series of clinically specific follow-up turns, evolve into something
functionally action-directing ("Based on everything you've told me, should I adjust my patient's
anticoagulation?"). Assessing each turn in isolation misses this cumulative dynamic entirely.
FDA should define the concept of an "interaction envelope" — the full scope of an intended-use
clinical interaction as defined by the manufacturer — and require that risk be assessed at the
interaction level, with the device's highest-risk plausible turn defining the interaction's risk
classification. Manufacturers must define and technically constrain the intended interaction
scope; any conversation that migrates outside the intended envelope should trigger a defined
response behavior (e.g., disclaimer, refusal, escalation).
Response to Question 6: Care Escalation — Asymmetric Harm Standards
FDA's attention to care escalation is well-placed. Under-escalation — failure to recognize that a
patient's condition warrants urgent intervention — represents the more severe harm in most
clinical contexts, with potential for direct patient injury or death. Over-escalation — directing
patients toward emergency care unnecessarily — creates system burden, patient anxiety, and
economic costs, but is rarely life-threatening. FDA should establish asymmetric harm standards
for escalation decisions: under-escalation threshold failures should be weighted more heavily in
risk classification and performance evaluation. However, over-escalation cannot be dismissed
entirely, particularly in resource-constrained environments where unnecessary escalations may
deprive other patients of timely care.
IV. SECTION V: PREMARKET EVALUATION — RESPONSES TO DISCUSSION QUESTIONS
7–17
Response to Question 7: The Competency-Based Approach Is Appropriate
Yes. The competency-based approach is the appropriate framework for evaluating generative AI
devices, and FDA deserves credit for developing it. It mirrors the structure by which human
clinical expertise is validated — board certification, specialty credentialing, simulation-based
assessment — and creates a principled basis for evaluating AI systems that operate similarly to
clinical judgment rather than to traditional deterministic software. WHM strongly supports this
framework as the foundation for FDA's regulatory approach.
Response to Question 8: Pathway Decision Matrix — A Prerequisite for Industry
Mapping the two-axis risk framework to specific regulatory pathways is not merely useful — it is
a prerequisite for the framework to function. Responsible manufacturers will not commit
development resources to a clinical AI program without knowing whether their device is headed
toward a 510(k) review, a De Novo classification request, or a PMA. These pathways carry
dramatically different time, cost, and evidence burdens. Investment decisions, partnership
structures, clinical study design, and commercial timelines all depend on this clarity.
FDA must publish a decision matrix — a clear, public reference linking risk tier (derived from the
two-axis framework) to regulatory pathway and corresponding premarket evidence
requirements — before finalizing this regulatory approach. Without it, the competency-based
framework, however conceptually sound, will produce prolonged pre-submission uncertainty for
manufacturers and an uneven submission landscape for FDA reviewers. We recommend that
the decision matrix be included as an appendix to any final guidance document, with explicit
examples for representative device categories.
Response to Question 9: Benchmarking Structure — Specialty Domain Modules Required
The benchmarking structure described in the Discussion Paper — organized across Safety (S.1–
S.3), Clinical Proficiency (E.1–E.4), Generalizability (R.1–R.2), and Agentic AI (A.1) — is a sound
general architecture. However, medical device generative AI does not exist in a domain-neutral
clinical environment. A cardiac AI device must be evaluated against cardiology-specific
benchmarks; a neuromodulation AI device requires benchmarks that reflect neuroscience
clinical context; a radiology AI tool requires imaging-specific performance standards.
FDA should establish specialty domain modules as required additions to the core benchmarking
structure for devices with defined clinical specialty applications. These modules should be
developed with input from relevant professional societies — the American College of
Cardiology, the American Academy of Neurology, the American College of Radiology, and their
counterparts — and should be updated periodically as clinical standards evolve. This will ensure
that benchmarking reflects the actual clinical context in which a device will operate, not a
generalized clinical abstraction.
Response to Question 10: Public Benchmark Contamination — A Serious Problem Requiring a
Structural Solution
Public benchmark contamination is not a theoretical concern. It is a documented phenomenon
in the AI industry: widely used public benchmarks have been incorporated — intentionally or
inadvertently — into training datasets, rendering them unreliable as independent performance
measures. For medical device applications, this is not merely an academic problem. A device
that appears to perform well on a contaminated benchmark may fail in clinical practice in ways
that could harm patients.
FDA should establish a sequestered national benchmark repository for medical AI, developed in
collaboration with NIST and maintained by an independent steward, analogous to NIST's existing
machine learning benchmarking challenges. Access to the sequestered evaluation set should be
tightly controlled, with submission and scoring handled through a blinded third-party process.
Sponsor-developed benchmarks should require third-party validation of construct validity
before FDA will accept them as primary evidence. The current reliance on publicly available
benchmarks, without structural controls against contamination, is a vulnerability that
adversarial actors could exploit and that well-intentioned manufacturers may inadvertently fall
into.
Response to Question 11: When Prospective Clinical Studies Are Required
The threshold for requiring a prospective clinical study should be calibrated to actual risk, not
applied uniformly across device categories. For Class II devices operating in HCP-supervised
environments, shadow deployment combined with retrospective clinical outcome evaluation
should generally satisfy the clinical confirmation requirement. Shadow deployment — in which
the AI device's outputs are recorded alongside clinical decisions made without reliance on those
outputs — provides meaningful real-world evidence without requiring the ethical and logistical
complexity of a prospective randomized design.
Prospective randomized clinical study design should be reserved for autonomous, patientfacing, Class III devices where AI-generated outputs will directly drive clinical actions without
contemporaneous HCP oversight. Applying prospective RCT requirements broadly would render
the competency-based framework economically unworkable for the vast majority of Class II
generative AI device categories and would create a competitive barrier that favors only the
largest manufacturers — precisely the outcome that least-burdensome principles are designed
to prevent.
Response to Questions 12 and 13: Synthetic Data — Caution and Clear Standards
Synthetic data has genuine utility in training and in augmenting performance evaluation for
underrepresented populations, but its limitations for validation are substantial and must be
explicitly addressed. Most critically: synthetic data generated by model families similar to the
device under evaluation creates circular validation. The model generates data that resembles its
own training distribution; evaluation on that synthetic data does not reveal the model's actual
failure modes.
FDA should require out-of-distribution (OOD) validation for any device that uses synthetic data
in its performance evaluation. OOD validation — testing on data that deliberately falls outside
the model's training distribution — is the most reliable way to detect blind spots that synthetic
in-distribution data would mask. Additionally, synthetic data is particularly problematic for
assessing performance across underrepresented racial, ethnic, and demographic subgroups,
precisely the populations for whom real-world generalizability gaps are most likely to emerge
and most likely to cause harm. For subpopulation performance validation, FDA should require
real-world data augmented with carefully documented synthetic sources, not synthetic data as a
primary evidentiary substitute.
Response to Question 14: Performance Comparators — Against the Right Standard
The choice of clinical performance comparator is consequential and must be handled carefully.
WHM recommends against using "median clinician in practice" as the primary performance
comparator for specialty AI devices. This standard, while operationally measurable, normalizes
the level of care that happens to prevail in clinical practice — which may reflect training gaps,
resource constraints, or systemic underperformance rather than optimal care. Clearing an AI
device because it performs at the level of the median clinician in practice does not ensure
patient benefit; it may simply replicate existing suboptimal performance at scale.
For specialty devices, the appropriate comparator is board-certified specialist performance
within the device's intended clinical domain and under conditions that reflect intended-use
clinical environments. For devices intended to expand access to specialist-level care in
underserved settings, comparison against delayed specialist review — reflecting the realistic
clinical alternative — is appropriate and clinically meaningful. FDA should codify these
comparator options in its decision matrix and guidance documents.
Response to Question 16: Third-Party Testing — Essential, But With Structural Safeguards
WHM strongly supports the use of independent third-party organizations in benchmarking and
clinical confirmation. Independent adjudication provides credibility, reduces sponsorship bias,
and improves the reliability of performance evidence. However, if FDA is not deliberate about
the structure of third-party testing, it risks creating a duopoly of approved testing organizations
— a small number of well-resourced labs with high barriers to entry that impose costs which
smaller manufacturers cannot sustain.
FDA should expand its Accreditation Scheme for Conformity Assessment (ASCA) program to
explicitly include qualified AI benchmarking organizations, with published, standardized fee
structures and published qualification criteria. Multiple qualified testing organizations should be
actively encouraged, including academic medical centers that have the domain expertise to
conduct specialty-relevant evaluation. The goal is a competitive, pluralistic testing ecosystem —
not a regulatory gatekeeping bottleneck.
Response to Question 17: Multimodal Architectures
Multimodal devices — those integrating imaging AI with clinical language model components,
for example — require explicit framework adaptation. A device that analyzes echocardiographic
imaging while simultaneously processing the patient's clinical history through a language model
presents a combined risk profile that cannot be adequately assessed by treating each modality
in isolation. Each modality must satisfy its domain-specific benchmarking requirements, and the
integrated device must demonstrate that combined performance is not materially degraded
relative to each component's standalone performance. FDA should issue a specific companion
guidance document for multimodal architectures, developed in consultation with imaging,
clinical AI, and specialty professional society stakeholders.
V. SECTION VI: POSTMARKET MONITORING — RESPONSES TO DISCUSSION
QUESTIONS 18–24
Response to Question 18: Accepting Greater Premarket Uncertainty — A Sound Trade
The proposal to accept greater premarket uncertainty in exchange for stronger, more rigorous
postmarket monitoring is conceptually sound and practically necessary for generative AI
devices. The nature of large language model and generative AI performance — inherently
probabilistic, contextually sensitive, and potentially subject to performance drift as clinical
language and practice norms evolve — makes it impossible to fully characterize safety and
efficacy through premarket evaluation alone. Real-world performance data, collected
systematically under defined conditions, will always provide the most meaningful signal.
WHM endorses this general principle with one critical condition: the postmarket monitoring
plan must be fully prespecified at the time of market authorization. It cannot be left to postclearance negotiation. The monitoring plan — including performance metrics, data collection
methodology, reporting cadence, triggering events for reassessment, and defined performance
thresholds that would trigger mandatory regulatory action — must be a legally enforceable
commitment, not an aspirational document. Manufacturers who commit to robust postmarket
monitoring should receive meaningful premarket flexibility; those who cannot commit to
rigorous monitoring should face correspondingly rigorous premarket standards.
Response to Question 19: Monitoring Cadence and Triggering Events
FDA should adopt a risk-tiered monitoring cadence as a default framework, subject to
modification based on device-specific characteristics:
• High-risk devices (action-taking, patient-facing, or involving irreversible outputs):
quarterly performance review with quarterly reporting to FDA.
• Medium-risk devices (action-directing, HCP-supervised): semi-annual performance
review with annual reporting to FDA, supplemented by event- triggered reporting as
defined below.
• Lower-risk devices (non-directive, decision-support): annual performance review with
annual reporting to FDA.
Triggering events requiring mandatory out-of-cycle reassessment should include, at minimum:
any material change to the underlying foundation model version or architecture; any MDR or
adverse event cluster meeting pre-specified thresholds; any payer policy change affecting the
device's indicated clinical use; any regulatory action by a peer regulatory authority (EU, Health
Canada, TGA) on the same or substantially similar device; and any published peer-reviewed
evidence that calls into question the clinical assumptions underlying the device's intended use.
Response to Question 20: Machine Supervisory Agents — Promising But Requires Standards
The use of machine-based supervisory agents to conduct real-time postmarket monitoring of
generative AI devices is promising and merits continued development. However, it creates a
recursive reliability problem that FDA must address directly: if AI is used to monitor AI, who
monitors the monitor?
Before FDA permits the use of machine supervisory agents as a primary postmarket monitoring
mechanism, it must establish performance standards for the supervisory agents themselves —
including validated sensitivity and specificity requirements for detecting performance
degradation, mandatory periodic evaluation of the supervisory agent's own accuracy, and a
mandatory human review layer triggered by supervisory agent alerts. FDA should initiate a
separate docket or technical working group to develop supervisory agent performance
standards in parallel with the primary GenAI device regulatory framework.
Response to Question 21: Stakeholder Accountability — Codify Roles to Prevent Diffusion
Postmarket monitoring is inherently a multi-stakeholder activity, but accountability diffusion is a
predictable consequence of undefined responsibilities. When everyone is theoretically
responsible, no one is operationally responsible.
FDA should codify the postmarket monitoring responsibilities of each stakeholder class in its
final guidance. Specifically: healthcare institutions — hospitals, integrated delivery networks,
ambulatory surgery centers — should bear responsibility for local workflow monitoring, userfacing adverse event detection, and institutional adverse event reporting. Professional societies
should develop and maintain domain-specific performance standards and should provide expert
review for MDR adjudication in their clinical domains. Foundation model developers should bear
responsibility for proactive disclosure of model changes and known failure modes. Device
manufacturers should retain primary reporting and remediation obligations. FDA should
establish clear reporting pathways for each stakeholder class, including electronic submission
standards and timeline requirements.
Response to Questions 22 and 23: PCCP Adaptation for Generative AI
Pre-Determined Change Control Plans must be substantially adapted for generative AI devices,
where the nature of anticipated changes — model updates, prompt modifications, retrieval
strategy refinements, guardrail adjustments — differs fundamentally from the firmware updates
or design modifications contemplated by existing PCCP guidance.
WHM recommends that FDA establish three explicit change categories for generative AI devices:
Category A — Locked Elements (require new premarket submission): Changes to training data
domain, core model architecture, intended use expansion beyond originally authorized scope,
and any change that materially alters the device's performance on its prespecified competency
benchmarks.
Category B — Supervised Elements (eligible for PCCP): Prompt engineering modifications,
retrieval-augmented generation strategy updates, output guardrail changes, and fine-tuning
within the original training data domain, provided that performance on competency
benchmarks is demonstrated to be maintained within pre-specified variation limits.
Category C — Monitored Elements (documentation only): User interface changes, response
formatting modifications, language localization updates, and administrative configuration
changes that do not affect clinical output.
FDA should publish explicit example lists for each category, updated annually as generative AI
development practices evolve. The current PCCP framework, applied without modification, will
create substantial regulatory uncertainty for manufacturers navigating the unique update
cadence of generative AI systems.
Response to Question 24: Third-Party Foundation Model Changes
Manufacturers bear responsibility for the performance of their devices, but they cannot fulfill
that responsibility if they lack notice of changes to the foundation models on which their devices
are built. FDA should require — as a condition of market authorization for devices built on thirdparty foundation models — that manufacturers demonstrate contractual notification rights
providing a minimum of ninety days advance notice of any material model version change,
including version deprecations. Shorter notice windows are inadequate for conducting
validation studies, executing change control procedures, and notifying FDA.
Technically, FDA should require manufacturers to implement model versioning pins with preproduction validation gates: no foundation model update should reach a deployed medical
device without first clearing a defined validation protocol. Automated deployment of foundation
model updates to live medical devices — without validation — should be explicitly prohibited.
For devices where contractual notification rights cannot be secured from the foundation model
developer, FDA should impose correspondingly higher premarket evidence requirements to
account for the elevated ongoing risk.
VI. SECTION VII: FOUNDATION MODEL MAFs AND AGENTIC AI — RESPONSES TO DISCUSSION
QUESTIONS 25–26
Response to Question 25: Foundation Model Device Master Files — Voluntary Is Insufficient
The concept of voluntary Foundation Model Device Master Files (MAFs) is structurally sound but
will not function as intended on a voluntary basis. Foundation model developers — large
technology companies and AI research organizations — have powerful competitive incentives to
withhold training data provenance, model architecture details, and known failure mode
documentation. Voluntary disclosure, in the absence of any regulatory consequence for nondisclosure, will produce a MAF system that the most responsible foundation model developers
participate in, while the least transparent actors opt out entirely. This creates a perverse
dynamic in which transparency is penalized.
WHM recommends that FDA establish a conditional market access mechanism: devices built on
foundation models for which no MAF has been filed, or for which MAF content falls below FDA's
minimum disclosure standards, should face a higher premarket evidence burden — specifically,
requiring more extensive sponsor- conducted benchmarking and third-party validation to
compensate for the lack of foundation model transparency. This creates a positive economic
incentive for foundation model developers to file complete MAFs without mandating disclosure
as a legal matter — a framework that is consistent with FDA's existing least- burdensome
principles while addressing the practical incentive problem.
When filed, MAFs should include, at minimum: a training data provenance summary describing
data sources, curation methodology, and known demographic representation gaps; a
documented inventory of known failure modes in clinical contexts, including clinical domains,
subpopulations, and task types where performance degradation has been observed;
performance benchmarks across demographic subgroups for relevant clinical tasks; a model
version history with meaningful changelog documentation; guardrail architecture description;
and an update notification protocol specifying how and when manufacturers using the
foundation model will be informed of material changes.
Response to Question 26: Agentic AI — The Most Significant Gap
Agentic AI represents the most significant gap in the Discussion Paper, and FDA is right to
identify it as requiring additional consideration. For agentic AI systems — those capable of
planning and executing multi-step actions in clinical environments with limited human oversight
— the current framework's concepts, developed primarily for advisory and action-directing
devices, are insufficient.
The stakes are highest for implantable and closed-loop devices. A neurostimulator that uses
LLM-analyzed sensor data to autonomously adjust therapy parameters in real time, a cardiac
rhythm management device that self-modifies pacing algorithms based on AI-generated clinical
pattern recognition, or a drug infusion system that titrates dosing based on agentic AI
interpretation of continuous monitoring data — each of these represents a risk profile that has
no adequate precedent in the current medical device regulatory framework. These are not
advisory tools. They are autonomous clinical actors.
FDA must address agentic AI across at least four specific domains:
First, human-in-the-loop requirements for irreversible actions must be mandatory. Any agentic
AI action that cannot be undone — device parameter changes that require a clinical procedure
to reverse, drug infusions above a defined dose threshold, surgical robot maneuvers — should
require a defined human confirmation step before execution. The Discussion Paper's discussion
of HCP oversight assumes advisory or action-directing devices; for agentic devices taking
irreversible actions, "oversight" must be redefined to mean pre-action authorization, not postaction review.
Second, FDA should develop the concept of an "action budget" for agentic AI systems — a
defined maximum number of consecutive autonomous actions a device may take without a
mandatory human confirmation checkpoint. Action budgets should be calibrated to device risk
tier, with higher-risk devices having more restrictive budgets, and should be a required element
of the device's approved intended use specification.
Third, fail-safe default behavior when supervisory connectivity is lost must be explicitly defined
and validated as part of premarket evaluation. An agentic AI device that loses its connection to
cloud-based supervisory infrastructure must have a validated, clinically safe default operating
mode that maintains patient safety without autonomous AI-directed action.
Fourth, agentic AI in surgical robotics requires entirely separate regulatory consideration. The
combination of physical action in an operating field, irreversibility, time pressure, and the
inherent complexity of surgical anatomy creates a risk profile that exceeds what the current
framework's concepts can adequately address. FDA should initiate a separate working group
specifically addressing agentic AI in surgical robotics, with participation from surgical
professional societies, medical device manufacturers, and patient safety organizations.
VII. ADDITIONAL RECOMMENDATIONS
A. Reimbursement and Coverage Alignment — A Formal FDA-CMS Initiative
FDA clearance creates no reimbursement rights. This is not a novel observation, but it has never
been more consequential than it will be for generative AI- enabled medical devices. CMS
currently operates without systematic mechanisms for assigning billing codes, conducting
coverage analysis, or establishing coverage with evidence development criteria for generative AI
devices. Commercial payers, who look to CMS for coverage signals, are similarly unprepared.
The practical consequence is predictable: FDA-cleared generative AI devices will be designated
"experimental and investigational" by payers for years after clearance, creating a market access
valley of death that delays patient access and destroys manufacturer commercial viability. This
outcome serves no one.
FDA should establish a formal liaison relationship with CMS's Coverage and Analysis Group
(CAG) to develop parallel evidentiary standards for clinical utility — standards that satisfy both
FDA's safety and effectiveness requirements and CMS's reasonable and necessary standard for
coverage. Ideally, clinical confirmation evidence generated for premarket evaluation should be
designed, from the outset, to satisfy both FDA and CMS evidentiary requirements. Coordinated
parallel review — not sequential, duplicative review — should be the structural goal. This would
represent a meaningful reduction in manufacturer burden and a material acceleration of patient
access to beneficial technology.
B. Labeling Requirements — Transparency as a Safety Mechanism
FDA should issue companion labeling guidance for generative AI-enabled devices that
establishes minimum labeling requirements for this device category. Required labeling elements
should include: clear disclosure that AI-generated content is presented to the user;
identification of the foundation model underlying the device, even if version-locked at a specific
release; performance limitations by demographic subpopulation, including any populations for
which validation data is limited; instructions for appropriate clinical oversight, including explicit
contraindications for unsupervised patient use where applicable; and a plain-language
description of the device's known failure modes and the circumstances under which clinician
judgment should take precedence.
Labeling transparency is a safety mechanism, not merely a disclosure formality. Clinicians who
understand a device's limitations are better positioned to use it appropriately; patients who
understand that AI-generated content may be imperfect are better positioned to raise concerns
when outputs seem inconsistent with their clinical experience.
C. Small Manufacturer Burden — A Dedicated Pathway Is Required
The competency-based premarket framework, as described, will impose costs that are not
uniformly distributed across manufacturers. Large, well-resourced manufacturers with
established regulatory teams, clinical research infrastructure, and third-party testing
relationships are comparatively well-positioned to navigate the framework. Small manufacturers
— startup medtech companies, academic spinouts, and specialty device developers — are not.
FDA should establish a Small Business GenAI Pathway that provides: fee waivers or reductions
for manufacturers below a defined revenue threshold (we recommend $10 million in projected
first-year revenue as a reasonable threshold); dedicated pre-submission consultation access
with GenAI-experienced reviewers, not general pre-submission staff; modular submission
formats that allow smaller manufacturers to build their applications in stages rather than
submitting a complete dossier; and public posting of anonymized pre-submission feedback to
build a shared knowledge base that reduces the cost of regulatory learning for all
manufacturers.
Small manufacturers are disproportionately responsible for genuinely novel clinical applications.
A regulatory framework that is navigable only by large incumbents is not innovation-enabling,
regardless of its technical sophistication.
D. International Harmonization — An Urgent Priority
FDA should treat international harmonization for generative AI medical devices as an urgent,
high-priority initiative. The EU AI Act and Medical Device Regulation together create a complex,
in some respects more prescriptive regulatory environment that global medtech manufacturers
must navigate in parallel with FDA requirements. Health Canada, Australia's TGA, and IMDRF are
each developing comparable frameworks. Without deliberate harmonization — beginning with
mutual recognition of benchmarking evidence and moving toward aligned clinical evidence
standards — global manufacturers will be required to conduct serial validation exercises across
jurisdictions, multiplying development costs and timelines.
FDA should immediately engage with EU notified bodies, Health Canada, TGA, and IMDRF
through existing international harmonization channels to develop a shared framework for
generative AI medical device evaluation, with a specific focus on benchmarking evidence mutual
recognition. The IMDRF work group structure is a natural vehicle for this coordination. A globally
harmonized generative AI medical device framework would represent a major advance for
patients and manufacturers worldwide and would appropriately position the United States as a
leader in beneficial innovation governance.
VIII. CONCLUSION
Walnut Hill Medical commends FDA for the quality of analysis reflected in this Discussion Paper
and for the deliberate, collaborative approach it represents. The competency-based framework
is conceptually right. The two-axis risk assessment is the appropriate organizing structure. The
engagement of stakeholders before finalizing policy reflects good administrative practice and
respect for the complexity of the challenges involved.
We urge FDA to take three specific structural commitments forward from this comment process.
First, publish a pathway decision matrix — linking risk tier to regulatory pathway to evidence
requirements — as a companion document to any final guidance, before that guidance takes
effect. Manufacturers cannot plan without it. Second, establish formal interagency coordination
with CMS to align clinical evidence standards with coverage and coding policy; FDA clearance
without coverage access serves neither patients nor innovation. Third, create a dedicated presubmission consultation program specifically for generative AI device sponsors, with particular
structural support for small and emerging manufacturers who represent the most dynamic
source of clinical innovation in this technology domain.
The regulatory framework FDA establishes for generative AI-enabled medical devices will shape
the trajectory of medical AI for a generation. Done well, it will enable a wave of genuinely
beneficial clinical technology to reach patients with appropriate safety assurance and
commercial viability. Done poorly, it will create barriers that drive innovation offshore or
underground. WHM is confident that FDA, with thoughtful stakeholder engagement and
structural commitment to least-burdensome principles, will get this right.
We thank FDA for the opportunity to comment and welcome the opportunity to discuss any
aspect of this submission in a public meeting or pre-submission consultation context.
Respectfully submitted,
Walnut Hill Medical Healthcare Reimbursement & Commercialization Strategy Consulting Dallas,
TX August 18, 2026
END OF COMMENT