← All 95 filings

Walnut Hill Medical

IndustryConsultantFiled August 18, 20266,118 words · 1 attachmentFDA-2026-N-7874-0003
“Payers will not reimburse what they cannot define.”

What they argued

RecovryAI’s one-line reading of the filing.

Q11 RCTs only for autonomous Class III; Q18 endorsed if monitoring plan prespecified; Q26 human confirmation before irreversible actions, action budgets.

Themes it raises

18 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Third, temporal urgency — whether the device operates in real-time versus asynchronous clinical workflows — creates meaningfully different risk profiles that the current framework does not distinguish.”
Whether the user can judge the outputFDA Q3, Q4
“Patients generally lack the clinical training to contextualize AI-generated outputs, assess limitations, or override recommendations that conflict with their broader clinical picture.”
Escalating too little and too muchFDA Q6
“FDA should establish asymmetric harm standards for escalation decisions: under-escalation threshold failures should be weighted more heavily in risk classification and performance evaluation.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“Mapping the two-axis risk framework to specific regulatory pathways is not merely useful — it is a prerequisite for the framework to function.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“FDA should establish a sequestered national benchmark repository for medical AI, developed in collaboration with NIST and maintained by an independent steward, analogous to NIST's existing machine learning benchmarking challenges.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“For Class II devices operating in HCP-supervised environments, shadow deployment combined with retrospective clinical outcome evaluation should generally satisfy the clinical confirmation requirement.”
Trading premarket certainty for postmarket monitoringFDA Q18
“WHM endorses this general principle with one critical condition: the postmarket monitoring plan must be fully prespecified at the time of market authorization.”
Watching the device after it shipsFDA Q19, Q20
“FDA should adopt a risk-tiered monitoring cadence as a default framework, subject to modification based on device-specific characteristics”
Who is accountable when something goes wrongFDA Q21
“FDA should codify the postmarket monitoring responsibilities of each stakeholder class in its final guidance.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“Automated deployment of foundation model updates to live medical devices — without validation — should be explicitly prohibited.”
Devices that plan and take actionsFDA Q26
“Any agentic AI action that cannot be undone — device parameter changes that require a clinical procedure to reverse, drug infusions above a defined dose threshold, surgical robot maneuvers — should require a defined human confirmation step before execution.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“for agentic devices taking irreversible actions, "oversight" must be redefined to mean pre-action authorization”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“Third, fail-safe default behavior when supervisory connectivity is lost must be explicitly defined and validated as part of premarket evaluation.”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“Additionally, synthetic data is particularly problematic for assessing performance across underrepresented racial, ethnic, and demographic subgroups, precisely the populations for whom real-world generalizability gaps are most likely to emerge and most likely to cause harm.”
How this fits rules that already existFDA Q8, Q9, Q16, Q25
“FDA should establish a formal liaison relationship with CMS's Coverage and Analysis Group (CAG) to develop parallel evidentiary standards for clinical utility — standards that satisfy both FDA's safety and effectiveness requirements and CMS's reasonable and necessary standard for coverage.”
What the rules cost sponsors and the marketNot asked by the FDA
“However, if FDA is not deliberate about the structure of third-party testing, it risks creating a duopoly of approved testing organizations — a small number of well-resourced labs with high barriers to entry that impose costs which smaller manufacturers cannot sustain.”
What patients are told and can demandFDA Q3
“Clinicians who understand a device's limitations are better positioned to use it appropriately; patients who understand that AI-generated content may be imperfect are better positioned to raise concerns when outputs seem inconsistent with their clinical experience.”
What counts as a reportable eventFDA Q19, Q20
“FDA should establish clear reporting pathways for each stakeholder class, including electronic submission standards and timeline requirements.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
Keep it, but add or change elements
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider the wording and specificity
Q3When clinical information goes straight to the patient, does the risk change, and what safeguards help without underestimating patients?
Treat patient-facing use differently
Base risk on the task and available safeguards
Q4Should it matter whether the clinician using the AI is a generalist or a specialist?
Assess the clinician’s task-specific knowledge
Test with the intended clinician group
Q5How is risk assessed when a conversation starts with non-directive information and drifts into action-directing?
Test whole conversations, not isolated answers
Enforce limits on what the conversation can do
Q6How should under-escalation be weighed against over-escalation?
Set stricter limits on dangerous missed escalations
Q7Is the two-step approach, benchmark the AI, then confirm it in clinical use, the right way to evaluate these devices?
Use benchmarking followed by clinical confirmation
Q8Should an AI device’s position on the risk map help decide how much evidence it must bring before market?
Link evidence requirements to the level of risk
Q9Do the ten benchmark competencies, from clinical knowledge to generalizability, add up to enough evidence of safety and effectiveness?
Use the structure, with additions or changes
Q10How can a benchmark score be shown to predict real-world behavior?
Protect test sets from exposure or contamination
Q11When can a device be confirmed without a prospective clinical study, and what earns that lighter path?
Some uses can be confirmed without a prospective study
Require prospective studies for specified higher-risk uses
Q14For open-ended AI outputs, who is the performance comparator: a clinician panel, generalists, specialists, or the human-AI team?
Judge against the applicable standard of care
Q16What role should independent third parties play?
Use independent parties to hold or maintain test assets
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Q17Does the approach still work for devices built on other model architectures, such as multimodal vision-language models and world models?
Adapt evaluation to the model architecture
Q18Can greater premarket uncertainty about a GenAI device’s benefit-risk profile be accepted through greater reliance on postmarket monitoring?
Allow it only under defined conditions
Q19How should an AI device be monitored after launch, and what sets the cadence?
Repeat performance testing on a schedule
Reassess after changes or safety signals
Q20Could AI supervisory agents help carry out postmarket monitoring?
Use AI monitoring with validated safeguards
Q21What roles should clinicians and institutions play in monitoring, without diluting manufacturer accountability?
Keep the manufacturer responsible for investigation and action
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Q24When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?
Identify and control the model version in use
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Q25Would voluntary Foundation Model Master Files be practical, and useful in premarket review?
Voluntary files alone are insufficient
Q26What extra oversight does an AI that plans and acts in multiple steps need?
Require human approval for specified consequential actions
Limit or test what the agent is allowed to do

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
Supports with conditions
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Acts
High-consequence work: Directs
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Walnut Hill Medical (WHM) is a healthcare reimbursement and commercialization strategy consulting firm with more than twenty years of engagement across medical device categories including cardiac, neuromodulation, vascular, implantable therapies, wearables, diagnostics, and surgical devices. We represent medical device manufacturers at every stage of commercialization.

We submit this comment to provide the perspective of medtech industry stakeholders on FDA’s proposed regulatory framework for generative AI-enabled medical devices. Our full comment is attached and addresses all 26 discussion questions posed in the Discussion Paper. Key positions include: (1) support for the competency-based evaluation framework; (2) the need for a published regulatory pathway decision matrix prior to framework finalization; (3) urgent interagency coordination between FDA and CMS on aligned evidentiary standards for coverage and reimbursement; and (4) practical modifications to the risk framework, benchmarking standards, PCCP change categories, and postmarket monitoring expectations. We respectfully urge FDA to prioritize regulatory clarity and least-burdensome principles as this framework develops.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

FORMAL PUBLIC COMMENT — FDA DOCKET NO. FDA-2026-N-7874
Submitted via: regulations.gov Date: August 18, 2026
To: FDA Dockets Management Staff Office of the Commissioner U.S. Food and Drug
Administration 5630 Fishers Lane, Room 1061 Rockville, MD 20852
Re: Docket No. FDA-2026-N-7874 — "Considerations for the Regulation of Generative AIEnabled Medical Devices: Discussion Paper and Request for Feedback"
From: Walnut Hill Medical (WHM) Healthcare Reimbursement & Commercialization Strategy
Consulting Dallas, TX channa@walnuthillmedical.com

I. INTRODUCTION
Walnut Hill Medical (WHM) is a healthcare reimbursement and commercialization strategy
consulting firm with more than twenty years of direct engagement across the full spectrum of
medical device categories — including cardiac devices, neuromodulation systems, vascular
therapies, implantable technologies, wearables, diagnostics, and surgical instruments. Our client
base spans early- stage medical device startups through established medtech companies, and
our work sits at the intersection of regulatory strategy, clinical reimbursement policy, payer
relations, coding and coverage analysis, and commercial implementation.
We submit this comment not as a regulatory law firm or academic institution, but as
practitioners who work daily alongside the manufacturers, clinical teams, hospital
administrators, and payer organizations that will be directly affected by FDA's regulatory
framework for generative AI-enabled medical devices. We have trained hundreds of medical
device representatives, clinical educators, prior authorization staff, billers and coders, and
hospital and ambulatory surgery center teams. We understand what happens when regulatory
frameworks create implementation friction — and when they enable innovation.
WHM's overall position is strongly supportive of FDA's collaborative approach in publishing this
Discussion Paper and requesting structured stakeholder input before finalizing regulatory policy.
The competency-based framework described in Sections V and VI represents a conceptually
sound and genuinely innovative approach. We urge FDA to continue in this spirit of
transparency, recognizing that generative AI-enabled medical devices represent a genuine
technological inflection point that demands new regulatory thinking rather than retrofitted
legacy frameworks.
Our central recommendation, developed throughout this comment, is that FDA must pair the
competency-based premarket framework with three structural commitments: (1) publication of
a clear pathway decision matrix before the framework is finalized, so manufacturers can make
rational development investments; (2) formal interagency coordination with CMS to align
clinical evidentiary standards with coverage and coding policy; and (3) a dedicated presubmission consultation mechanism for generative AI device sponsors, with particular
accommodation for small and emerging manufacturers.
II. GENERAL OBSERVATIONS
A. Competency-Based Evaluation: Conceptually Right
The Discussion Paper's central innovation — evaluating generative AI devices against defined
competencies rather than fixed performance thresholds — is analytically sound and reflects how
clinical expertise itself is validated. Just as a board-certified cardiologist is assessed not by a
single test score but by demonstrated competency across a spectrum of clinical scenarios, a
generative AI-enabled diagnostic tool should be evaluated against structured performance
domains that reflect the full range of clinical contexts in which it will operate. WHM strongly
endorses this organizing principle and encourages FDA to develop it with precision and rigor.

B. The Critical Gap: Reimbursement and Coverage Alignment
The most significant structural omission in the Discussion Paper is the absence of any discussion
of reimbursement and coverage alignment. FDA clearance or approval does not create
reimbursement rights. CMS currently operates without a systematic approach to coding and
coverage for generative AI-enabled medical devices, and commercial payers follow CMS signals
closely. If FDA clears a generative AI device without parallel work at CMS to establish a coding
pathway, coverage criteria, and evidentiary standards for coverage with evidence development,
the cleared device will face a market access valley of death that no manufacturer — and no
patient — benefits from.
Payers will not reimburse what they cannot define. They will not cover devices for which no
billing code exists, and they will apply stringent "experimental and investigational" designations
to generative AI tools until clinical evidence standards are harmonized between FDA and CMS.
The competency-based evidence FDA requires for clearance should map directly — and by
design — to the clinical evidence CMS and commercial payers require for coverage decisions.
This alignment does not happen automatically. It requires formal interagency coordination, and
FDA should commit to establishing it as part of this regulatory initiative.

C. International Harmonization Gaps
FDA's framework will not exist in isolation. The European Union's AI Act and the Medical Device
Regulation together create a parallel, in some respects more prescriptive, regulatory
environment for AI-enabled medical devices. Health Canada, Australia's TGA, and IMDRF are
each developing or refining comparable frameworks. Without deliberate harmonization — at
minimum, mutual recognition of benchmarking evidence — global medtech manufacturers will
face serial validation exercises across jurisdictions, multiplying the development burden and
creating significant competitive disadvantage relative to manufacturers operating in less
rigorous markets. FDA should treat international harmonization as a first-order priority, not an
afterthought.

D. A Tiered Pathway Map Is Essential
The Discussion Paper describes regulatory concepts without providing a clear map of which
generative AI device type belongs in which regulatory pathway. Before manufacturers can make
rational development investments, FDA must publish a decision matrix that links device risk tier
(derived from the two-axis framework) to the applicable regulatory pathway — 510(k), De Novo,
or PMA — and to the corresponding premarket evidence requirements. Absent this clarity,
responsible manufacturers will over-invest in evidence generation out of regulatory uncertainty,
and less responsible actors will under-invest and submit inadequate applications, taxing FDA
review resources.

III. SECTION IV: RISK ASSESSMENT — RESPONSES TO DISCUSSION QUESTIONS 1–6
The two-axis framework — characterizing device activity on a spectrum from non- directive to
action-directing to action-taking, crossed with an assessment of the consequences of incorrect
output — is the right organizing principle for generative AI risk stratification. WHM supports it
strongly and offers the following refinements.
Response to Question 1: Additional Dimensions of Risk
The two-axis framework captures the core of the risk landscape but requires three additional
dimensions to be complete.
First, third-party model dependency is a material risk factor not reflected in the current
framework. A generative AI device built entirely on a proprietary, auditable, manufacturercontrolled model presents a fundamentally different risk profile than a device built on an
opaque third-party foundation model over which the manufacturer has no visibility, no
contractual governance rights, and no notification of changes. The current framework treats
these devices identically. They are not identical. FDA should add third-party model dependency
as an explicit risk dimension, with devices heavily dependent on non-disclosed, externally
controlled foundation models receiving higher baseline risk classification.
Second, reversibility must be treated as a first-order risk factor. An AI- generated
recommendation that a clinician reviews and acts upon over the course of hours is materially
different from an AI-generated command that triggers an immediate, irreversible action — a
drug infusion, a neurostimulation parameter change, a surgical robot maneuver. The current
framework's consequence axis captures severity but does not fully capture irreversibility. A
separate reversibility dimension would strengthen the framework's clinical fidelity.
Third, temporal urgency — whether the device operates in real-time versus asynchronous
clinical workflows — creates meaningfully different risk profiles that the current framework
does not distinguish.
Real-time autonomous action under time pressure (e.g., AI-guided
emergency drug dosing) warrants higher scrutiny than asynchronous AI-generated care plan
recommendations reviewed by a clinician before acting.
Response to Question 2: Directiveness as a Spectrum — The Need for Safe Harbor
WHM supports characterizing directiveness as a spectrum rather than a binary, but
manufacturers require a safe harbor bright line to plan development programs. Without at least
one clear presumptive standard, manufacturers face perpetual uncertainty about where their
device falls on the continuum. We recommend that FDA establish a rebuttable presumption: any
generative AI output that includes specific numeric parameters — drug doses, device therapy
thresholds, surgical dimensions — is presumptively action-directing, regardless of framing
language. Manufacturers who believe their device should be characterized differently can rebut
the presumption through pre-submission consultation. This approach provides clarity while
preserving flexibility.
Response to Question 3: Patient-Facing vs. HCP-Facing Applications
FDA is right to treat patient-facing and HCP-facing applications differently. Patients generally
lack the clinical training to contextualize AI-generated outputs, assess limitations, or override
recommendations that conflict with their broader clinical picture.
WHM supports differential
regulatory treatment.
However, FDA must be careful not to overreach. Patient-facing informational tools — symptom
checkers, post-discharge medication reminders, general wellness education — should not be
regulated as action-directing devices merely because a patient reads and acts upon them. The
current framing of "action-taking" risk in patient-facing contexts risks creating a category of
regulation so broad that it encompasses general health information technology entirely. FDA
should define clear boundaries: patient-facing tools that include specific clinical
recommendations (e.g., "your symptoms indicate you should go to the emergency room") are
meaningfully different from those that provide general information without clinical specificity
(e.g., "these are common side effects of your medication"). The former warrants closer scrutiny;
the latter generally should not.
Response to Question 4: Generalist vs. Specialist HCP-Facing Applications
The distinction between generalist and specialist HCP-facing applications is clinically significant.
A generative AI tool designed for use by a board- certified interventional cardiologist carries
different risk implications than the same tool used by a primary care physician without specialty
training in interpreting its outputs. FDA should address this distinction explicitly in its labeling
and intended use requirements, requiring manufacturers to specify the intended clinical user's
competency level and to validate performance within that specific user population.
Response to Question 5: Multi-Turn Conversations — Assess at the Interaction Level
Risk in multi-turn conversational AI cannot be assessed at the level of individual turns. A single
exchange that begins with a non-directive inquiry ("What are the risk factors for atrial
fibrillation?") may, through a series of clinically specific follow-up turns, evolve into something
functionally action-directing ("Based on everything you've told me, should I adjust my patient's
anticoagulation?"). Assessing each turn in isolation misses this cumulative dynamic entirely.
FDA should define the concept of an "interaction envelope" — the full scope of an intended-use
clinical interaction as defined by the manufacturer — and require that risk be assessed at the
interaction level, with the device's highest-risk plausible turn defining the interaction's risk
classification. Manufacturers must define and technically constrain the intended interaction
scope; any conversation that migrates outside the intended envelope should trigger a defined
response behavior (e.g., disclaimer, refusal, escalation).
Response to Question 6: Care Escalation — Asymmetric Harm Standards
FDA's attention to care escalation is well-placed. Under-escalation — failure to recognize that a
patient's condition warrants urgent intervention — represents the more severe harm in most
clinical contexts, with potential for direct patient injury or death. Over-escalation — directing
patients toward emergency care unnecessarily — creates system burden, patient anxiety, and
economic costs, but is rarely life-threatening. FDA should establish asymmetric harm standards
for escalation decisions: under-escalation threshold failures should be weighted more heavily in
risk classification and performance evaluation.
However, over-escalation cannot be dismissed
entirely, particularly in resource-constrained environments where unnecessary escalations may
deprive other patients of timely care.

IV. SECTION V: PREMARKET EVALUATION — RESPONSES TO DISCUSSION QUESTIONS
7–17
Response to Question 7: The Competency-Based Approach Is Appropriate
Yes. The competency-based approach is the appropriate framework for evaluating generative AI
devices, and FDA deserves credit for developing it. It mirrors the structure by which human
clinical expertise is validated — board certification, specialty credentialing, simulation-based
assessment — and creates a principled basis for evaluating AI systems that operate similarly to
clinical judgment rather than to traditional deterministic software. WHM strongly supports this
framework as the foundation for FDA's regulatory approach.
Response to Question 8: Pathway Decision Matrix — A Prerequisite for Industry
Mapping the two-axis risk framework to specific regulatory pathways is not merely useful — it is
a prerequisite for the framework to function.
Responsible manufacturers will not commit
development resources to a clinical AI program without knowing whether their device is headed
toward a 510(k) review, a De Novo classification request, or a PMA. These pathways carry
dramatically different time, cost, and evidence burdens. Investment decisions, partnership
structures, clinical study design, and commercial timelines all depend on this clarity.
FDA must publish a decision matrix — a clear, public reference linking risk tier (derived from the
two-axis framework) to regulatory pathway and corresponding premarket evidence
requirements — before finalizing this regulatory approach. Without it, the competency-based
framework, however conceptually sound, will produce prolonged pre-submission uncertainty for
manufacturers and an uneven submission landscape for FDA reviewers. We recommend that
the decision matrix be included as an appendix to any final guidance document, with explicit
examples for representative device categories.
Response to Question 9: Benchmarking Structure — Specialty Domain Modules Required
The benchmarking structure described in the Discussion Paper — organized across Safety (S.1–
S.3), Clinical Proficiency (E.1–E.4), Generalizability (R.1–R.2), and Agentic AI (A.1) — is a sound
general architecture. However, medical device generative AI does not exist in a domain-neutral
clinical environment. A cardiac AI device must be evaluated against cardiology-specific
benchmarks; a neuromodulation AI device requires benchmarks that reflect neuroscience
clinical context; a radiology AI tool requires imaging-specific performance standards.
FDA should establish specialty domain modules as required additions to the core benchmarking
structure for devices with defined clinical specialty applications. These modules should be
developed with input from relevant professional societies — the American College of
Cardiology, the American Academy of Neurology, the American College of Radiology, and their
counterparts — and should be updated periodically as clinical standards evolve. This will ensure
that benchmarking reflects the actual clinical context in which a device will operate, not a
generalized clinical abstraction.
Response to Question 10: Public Benchmark Contamination — A Serious Problem Requiring a
Structural Solution
Public benchmark contamination is not a theoretical concern. It is a documented phenomenon
in the AI industry: widely used public benchmarks have been incorporated — intentionally or
inadvertently — into training datasets, rendering them unreliable as independent performance
measures. For medical device applications, this is not merely an academic problem. A device
that appears to perform well on a contaminated benchmark may fail in clinical practice in ways
that could harm patients.
FDA should establish a sequestered national benchmark repository for medical AI, developed in
collaboration with NIST and maintained by an independent steward, analogous to NIST's existing
machine learning benchmarking challenges.
Access to the sequestered evaluation set should be
tightly controlled, with submission and scoring handled through a blinded third-party process.
Sponsor-developed benchmarks should require third-party validation of construct validity
before FDA will accept them as primary evidence. The current reliance on publicly available
benchmarks, without structural controls against contamination, is a vulnerability that
adversarial actors could exploit and that well-intentioned manufacturers may inadvertently fall
into.
Response to Question 11: When Prospective Clinical Studies Are Required
The threshold for requiring a prospective clinical study should be calibrated to actual risk, not
applied uniformly across device categories. For Class II devices operating in HCP-supervised
environments, shadow deployment combined with retrospective clinical outcome evaluation
should generally satisfy the clinical confirmation requirement.
Shadow deployment — in which
the AI device's outputs are recorded alongside clinical decisions made without reliance on those
outputs — provides meaningful real-world evidence without requiring the ethical and logistical
complexity of a prospective randomized design.
Prospective randomized clinical study design should be reserved for autonomous, patientfacing, Class III devices where AI-generated outputs will directly drive clinical actions without
contemporaneous HCP oversight. Applying prospective RCT requirements broadly would render
the competency-based framework economically unworkable for the vast majority of Class II
generative AI device categories and would create a competitive barrier that favors only the
largest manufacturers — precisely the outcome that least-burdensome principles are designed
to prevent.
Response to Questions 12 and 13: Synthetic Data — Caution and Clear Standards
Synthetic data has genuine utility in training and in augmenting performance evaluation for
underrepresented populations, but its limitations for validation are substantial and must be
explicitly addressed. Most critically: synthetic data generated by model families similar to the
device under evaluation creates circular validation. The model generates data that resembles its
own training distribution; evaluation on that synthetic data does not reveal the model's actual
failure modes.
FDA should require out-of-distribution (OOD) validation for any device that uses synthetic data
in its performance evaluation. OOD validation — testing on data that deliberately falls outside
the model's training distribution — is the most reliable way to detect blind spots that synthetic
in-distribution data would mask. Additionally, synthetic data is particularly problematic for
assessing performance across underrepresented racial, ethnic, and demographic subgroups,
precisely the populations for whom real-world generalizability gaps are most likely to emerge
and most likely to cause harm.
For subpopulation performance validation, FDA should require
real-world data augmented with carefully documented synthetic sources, not synthetic data as a
primary evidentiary substitute.
Response to Question 14: Performance Comparators — Against the Right Standard
The choice of clinical performance comparator is consequential and must be handled carefully.
WHM recommends against using "median clinician in practice" as the primary performance
comparator for specialty AI devices. This standard, while operationally measurable, normalizes
the level of care that happens to prevail in clinical practice — which may reflect training gaps,
resource constraints, or systemic underperformance rather than optimal care. Clearing an AI
device because it performs at the level of the median clinician in practice does not ensure
patient benefit; it may simply replicate existing suboptimal performance at scale.
For specialty devices, the appropriate comparator is board-certified specialist performance
within the device's intended clinical domain and under conditions that reflect intended-use
clinical environments. For devices intended to expand access to specialist-level care in
underserved settings, comparison against delayed specialist review — reflecting the realistic
clinical alternative — is appropriate and clinically meaningful. FDA should codify these
comparator options in its decision matrix and guidance documents.
Response to Question 16: Third-Party Testing — Essential, But With Structural Safeguards
WHM strongly supports the use of independent third-party organizations in benchmarking and
clinical confirmation. Independent adjudication provides credibility, reduces sponsorship bias,
and improves the reliability of performance evidence. However, if FDA is not deliberate about
the structure of third-party testing, it risks creating a duopoly of approved testing organizations
— a small number of well-resourced labs with high barriers to entry that impose costs which
smaller manufacturers cannot sustain.

FDA should expand its Accreditation Scheme for Conformity Assessment (ASCA) program to
explicitly include qualified AI benchmarking organizations, with published, standardized fee
structures and published qualification criteria. Multiple qualified testing organizations should be
actively encouraged, including academic medical centers that have the domain expertise to
conduct specialty-relevant evaluation. The goal is a competitive, pluralistic testing ecosystem —
not a regulatory gatekeeping bottleneck.
Response to Question 17: Multimodal Architectures
Multimodal devices — those integrating imaging AI with clinical language model components,
for example — require explicit framework adaptation. A device that analyzes echocardiographic
imaging while simultaneously processing the patient's clinical history through a language model
presents a combined risk profile that cannot be adequately assessed by treating each modality
in isolation. Each modality must satisfy its domain-specific benchmarking requirements, and the
integrated device must demonstrate that combined performance is not materially degraded
relative to each component's standalone performance. FDA should issue a specific companion
guidance document for multimodal architectures, developed in consultation with imaging,
clinical AI, and specialty professional society stakeholders.

V. SECTION VI: POSTMARKET MONITORING — RESPONSES TO DISCUSSION
QUESTIONS 18–24
Response to Question 18: Accepting Greater Premarket Uncertainty — A Sound Trade
The proposal to accept greater premarket uncertainty in exchange for stronger, more rigorous
postmarket monitoring is conceptually sound and practically necessary for generative AI
devices. The nature of large language model and generative AI performance — inherently
probabilistic, contextually sensitive, and potentially subject to performance drift as clinical
language and practice norms evolve — makes it impossible to fully characterize safety and
efficacy through premarket evaluation alone. Real-world performance data, collected
systematically under defined conditions, will always provide the most meaningful signal.
WHM endorses this general principle with one critical condition: the postmarket monitoring
plan must be fully prespecified at the time of market authorization.
It cannot be left to postclearance negotiation. The monitoring plan — including performance metrics, data collection
methodology, reporting cadence, triggering events for reassessment, and defined performance
thresholds that would trigger mandatory regulatory action — must be a legally enforceable
commitment, not an aspirational document. Manufacturers who commit to robust postmarket
monitoring should receive meaningful premarket flexibility; those who cannot commit to
rigorous monitoring should face correspondingly rigorous premarket standards.
Response to Question 19: Monitoring Cadence and Triggering Events
FDA should adopt a risk-tiered monitoring cadence as a default framework, subject to
modification based on device-specific characteristics
:

• High-risk devices (action-taking, patient-facing, or involving irreversible outputs):
quarterly performance review with quarterly reporting to FDA.
• Medium-risk devices (action-directing, HCP-supervised): semi-annual performance
review with annual reporting to FDA, supplemented by event- triggered reporting as
defined below.
• Lower-risk devices (non-directive, decision-support): annual performance review with
annual reporting to FDA.
Triggering events requiring mandatory out-of-cycle reassessment should include, at minimum:
any material change to the underlying foundation model version or architecture; any MDR or
adverse event cluster meeting pre-specified thresholds; any payer policy change affecting the
device's indicated clinical use; any regulatory action by a peer regulatory authority (EU, Health
Canada, TGA) on the same or substantially similar device; and any published peer-reviewed
evidence that calls into question the clinical assumptions underlying the device's intended use.
Response to Question 20: Machine Supervisory Agents — Promising But Requires Standards
The use of machine-based supervisory agents to conduct real-time postmarket monitoring of
generative AI devices is promising and merits continued development. However, it creates a
recursive reliability problem that FDA must address directly: if AI is used to monitor AI, who
monitors the monitor?
Before FDA permits the use of machine supervisory agents as a primary postmarket monitoring
mechanism, it must establish performance standards for the supervisory agents themselves —
including validated sensitivity and specificity requirements for detecting performance
degradation, mandatory periodic evaluation of the supervisory agent's own accuracy, and a
mandatory human review layer triggered by supervisory agent alerts. FDA should initiate a
separate docket or technical working group to develop supervisory agent performance
standards in parallel with the primary GenAI device regulatory framework.
Response to Question 21: Stakeholder Accountability — Codify Roles to Prevent Diffusion
Postmarket monitoring is inherently a multi-stakeholder activity, but accountability diffusion is a
predictable consequence of undefined responsibilities. When everyone is theoretically
responsible, no one is operationally responsible.
FDA should codify the postmarket monitoring responsibilities of each stakeholder class in its
final guidance.
Specifically: healthcare institutions — hospitals, integrated delivery networks,
ambulatory surgery centers — should bear responsibility for local workflow monitoring, userfacing adverse event detection, and institutional adverse event reporting. Professional societies
should develop and maintain domain-specific performance standards and should provide expert
review for MDR adjudication in their clinical domains. Foundation model developers should bear
responsibility for proactive disclosure of model changes and known failure modes. Device
manufacturers should retain primary reporting and remediation obligations. FDA should
establish clear reporting pathways for each stakeholder class, including electronic submission
standards and timeline requirements.

Response to Questions 22 and 23: PCCP Adaptation for Generative AI
Pre-Determined Change Control Plans must be substantially adapted for generative AI devices,
where the nature of anticipated changes — model updates, prompt modifications, retrieval
strategy refinements, guardrail adjustments — differs fundamentally from the firmware updates
or design modifications contemplated by existing PCCP guidance.
WHM recommends that FDA establish three explicit change categories for generative AI devices:
Category A — Locked Elements (require new premarket submission): Changes to training data
domain, core model architecture, intended use expansion beyond originally authorized scope,
and any change that materially alters the device's performance on its prespecified competency
benchmarks.
Category B — Supervised Elements (eligible for PCCP): Prompt engineering modifications,
retrieval-augmented generation strategy updates, output guardrail changes, and fine-tuning
within the original training data domain, provided that performance on competency
benchmarks is demonstrated to be maintained within pre-specified variation limits.
Category C — Monitored Elements (documentation only): User interface changes, response
formatting modifications, language localization updates, and administrative configuration
changes that do not affect clinical output.
FDA should publish explicit example lists for each category, updated annually as generative AI
development practices evolve. The current PCCP framework, applied without modification, will
create substantial regulatory uncertainty for manufacturers navigating the unique update
cadence of generative AI systems.
Response to Question 24: Third-Party Foundation Model Changes
Manufacturers bear responsibility for the performance of their devices, but they cannot fulfill
that responsibility if they lack notice of changes to the foundation models on which their devices
are built. FDA should require — as a condition of market authorization for devices built on thirdparty foundation models — that manufacturers demonstrate contractual notification rights
providing a minimum of ninety days advance notice of any material model version change,
including version deprecations. Shorter notice windows are inadequate for conducting
validation studies, executing change control procedures, and notifying FDA.
Technically, FDA should require manufacturers to implement model versioning pins with preproduction validation gates: no foundation model update should reach a deployed medical
device without first clearing a defined validation protocol. Automated deployment of foundation
model updates to live medical devices — without validation — should be explicitly prohibited.

For devices where contractual notification rights cannot be secured from the foundation model
developer, FDA should impose correspondingly higher premarket evidence requirements to
account for the elevated ongoing risk.
VI. SECTION VII: FOUNDATION MODEL MAFs AND AGENTIC AI — RESPONSES TO DISCUSSION
QUESTIONS 25–26
Response to Question 25: Foundation Model Device Master Files — Voluntary Is Insufficient
The concept of voluntary Foundation Model Device Master Files (MAFs) is structurally sound but
will not function as intended on a voluntary basis. Foundation model developers — large
technology companies and AI research organizations — have powerful competitive incentives to
withhold training data provenance, model architecture details, and known failure mode
documentation. Voluntary disclosure, in the absence of any regulatory consequence for nondisclosure, will produce a MAF system that the most responsible foundation model developers
participate in, while the least transparent actors opt out entirely. This creates a perverse
dynamic in which transparency is penalized.
WHM recommends that FDA establish a conditional market access mechanism: devices built on
foundation models for which no MAF has been filed, or for which MAF content falls below FDA's
minimum disclosure standards, should face a higher premarket evidence burden — specifically,
requiring more extensive sponsor- conducted benchmarking and third-party validation to
compensate for the lack of foundation model transparency. This creates a positive economic
incentive for foundation model developers to file complete MAFs without mandating disclosure
as a legal matter — a framework that is consistent with FDA's existing least- burdensome
principles while addressing the practical incentive problem.
When filed, MAFs should include, at minimum: a training data provenance summary describing
data sources, curation methodology, and known demographic representation gaps; a
documented inventory of known failure modes in clinical contexts, including clinical domains,
subpopulations, and task types where performance degradation has been observed;
performance benchmarks across demographic subgroups for relevant clinical tasks; a model
version history with meaningful changelog documentation; guardrail architecture description;
and an update notification protocol specifying how and when manufacturers using the
foundation model will be informed of material changes.
Response to Question 26: Agentic AI — The Most Significant Gap
Agentic AI represents the most significant gap in the Discussion Paper, and FDA is right to
identify it as requiring additional consideration. For agentic AI systems — those capable of
planning and executing multi-step actions in clinical environments with limited human oversight
— the current framework's concepts, developed primarily for advisory and action-directing
devices, are insufficient.
The stakes are highest for implantable and closed-loop devices. A neurostimulator that uses
LLM-analyzed sensor data to autonomously adjust therapy parameters in real time, a cardiac
rhythm management device that self-modifies pacing algorithms based on AI-generated clinical
pattern recognition, or a drug infusion system that titrates dosing based on agentic AI
interpretation of continuous monitoring data — each of these represents a risk profile that has
no adequate precedent in the current medical device regulatory framework. These are not
advisory tools. They are autonomous clinical actors.
FDA must address agentic AI across at least four specific domains:
First, human-in-the-loop requirements for irreversible actions must be mandatory. Any agentic
AI action that cannot be undone — device parameter changes that require a clinical procedure
to reverse, drug infusions above a defined dose threshold, surgical robot maneuvers — should
require a defined human confirmation step before execution.
The Discussion Paper's discussion
of HCP oversight assumes advisory or action-directing devices; for agentic devices taking
irreversible actions, "oversight" must be redefined to mean pre-action authorization
, not postaction review.
Second, FDA should develop the concept of an "action budget" for agentic AI systems — a
defined maximum number of consecutive autonomous actions a device may take without a
mandatory human confirmation checkpoint. Action budgets should be calibrated to device risk
tier, with higher-risk devices having more restrictive budgets, and should be a required element
of the device's approved intended use specification.
Third, fail-safe default behavior when supervisory connectivity is lost must be explicitly defined
and validated as part of premarket evaluation.
An agentic AI device that loses its connection to
cloud-based supervisory infrastructure must have a validated, clinically safe default operating
mode that maintains patient safety without autonomous AI-directed action.
Fourth, agentic AI in surgical robotics requires entirely separate regulatory consideration. The
combination of physical action in an operating field, irreversibility, time pressure, and the
inherent complexity of surgical anatomy creates a risk profile that exceeds what the current
framework's concepts can adequately address. FDA should initiate a separate working group
specifically addressing agentic AI in surgical robotics, with participation from surgical
professional societies, medical device manufacturers, and patient safety organizations.

VII. ADDITIONAL RECOMMENDATIONS
A. Reimbursement and Coverage Alignment — A Formal FDA-CMS Initiative
FDA clearance creates no reimbursement rights. This is not a novel observation, but it has never
been more consequential than it will be for generative AI- enabled medical devices. CMS
currently operates without systematic mechanisms for assigning billing codes, conducting
coverage analysis, or establishing coverage with evidence development criteria for generative AI
devices. Commercial payers, who look to CMS for coverage signals, are similarly unprepared.
The practical consequence is predictable: FDA-cleared generative AI devices will be designated
"experimental and investigational" by payers for years after clearance, creating a market access
valley of death that delays patient access and destroys manufacturer commercial viability. This
outcome serves no one.
FDA should establish a formal liaison relationship with CMS's Coverage and Analysis Group
(CAG) to develop parallel evidentiary standards for clinical utility — standards that satisfy both
FDA's safety and effectiveness requirements and CMS's reasonable and necessary standard for
coverage.
Ideally, clinical confirmation evidence generated for premarket evaluation should be
designed, from the outset, to satisfy both FDA and CMS evidentiary requirements. Coordinated
parallel review — not sequential, duplicative review — should be the structural goal. This would
represent a meaningful reduction in manufacturer burden and a material acceleration of patient
access to beneficial technology.

B. Labeling Requirements — Transparency as a Safety Mechanism
FDA should issue companion labeling guidance for generative AI-enabled devices that
establishes minimum labeling requirements for this device category. Required labeling elements
should include: clear disclosure that AI-generated content is presented to the user;
identification of the foundation model underlying the device, even if version-locked at a specific
release; performance limitations by demographic subpopulation, including any populations for
which validation data is limited; instructions for appropriate clinical oversight, including explicit
contraindications for unsupervised patient use where applicable; and a plain-language
description of the device's known failure modes and the circumstances under which clinician
judgment should take precedence.
Labeling transparency is a safety mechanism, not merely a disclosure formality. Clinicians who
understand a device's limitations are better positioned to use it appropriately; patients who
understand that AI-generated content may be imperfect are better positioned to raise concerns
when outputs seem inconsistent with their clinical experience.

C. Small Manufacturer Burden — A Dedicated Pathway Is Required
The competency-based premarket framework, as described, will impose costs that are not
uniformly distributed across manufacturers. Large, well-resourced manufacturers with
established regulatory teams, clinical research infrastructure, and third-party testing
relationships are comparatively well-positioned to navigate the framework. Small manufacturers
— startup medtech companies, academic spinouts, and specialty device developers — are not.
FDA should establish a Small Business GenAI Pathway that provides: fee waivers or reductions
for manufacturers below a defined revenue threshold (we recommend $10 million in projected
first-year revenue as a reasonable threshold); dedicated pre-submission consultation access
with GenAI-experienced reviewers, not general pre-submission staff; modular submission
formats that allow smaller manufacturers to build their applications in stages rather than
submitting a complete dossier; and public posting of anonymized pre-submission feedback to
build a shared knowledge base that reduces the cost of regulatory learning for all
manufacturers.
Small manufacturers are disproportionately responsible for genuinely novel clinical applications.
A regulatory framework that is navigable only by large incumbents is not innovation-enabling,
regardless of its technical sophistication.

D. International Harmonization — An Urgent Priority
FDA should treat international harmonization for generative AI medical devices as an urgent,
high-priority initiative. The EU AI Act and Medical Device Regulation together create a complex,
in some respects more prescriptive regulatory environment that global medtech manufacturers
must navigate in parallel with FDA requirements. Health Canada, Australia's TGA, and IMDRF are
each developing comparable frameworks. Without deliberate harmonization — beginning with
mutual recognition of benchmarking evidence and moving toward aligned clinical evidence
standards — global manufacturers will be required to conduct serial validation exercises across
jurisdictions, multiplying development costs and timelines.
FDA should immediately engage with EU notified bodies, Health Canada, TGA, and IMDRF
through existing international harmonization channels to develop a shared framework for
generative AI medical device evaluation, with a specific focus on benchmarking evidence mutual
recognition. The IMDRF work group structure is a natural vehicle for this coordination. A globally
harmonized generative AI medical device framework would represent a major advance for
patients and manufacturers worldwide and would appropriately position the United States as a
leader in beneficial innovation governance.

VIII. CONCLUSION
Walnut Hill Medical commends FDA for the quality of analysis reflected in this Discussion Paper
and for the deliberate, collaborative approach it represents. The competency-based framework
is conceptually right. The two-axis risk assessment is the appropriate organizing structure. The
engagement of stakeholders before finalizing policy reflects good administrative practice and
respect for the complexity of the challenges involved.
We urge FDA to take three specific structural commitments forward from this comment process.
First, publish a pathway decision matrix — linking risk tier to regulatory pathway to evidence
requirements — as a companion document to any final guidance, before that guidance takes
effect. Manufacturers cannot plan without it. Second, establish formal interagency coordination
with CMS to align clinical evidence standards with coverage and coding policy; FDA clearance
without coverage access serves neither patients nor innovation. Third, create a dedicated presubmission consultation program specifically for generative AI device sponsors, with particular
structural support for small and emerging manufacturers who represent the most dynamic
source of clinical innovation in this technology domain.
The regulatory framework FDA establishes for generative AI-enabled medical devices will shape
the trajectory of medical AI for a generation. Done well, it will enable a wave of genuinely
beneficial clinical technology to reach patients with appropriate safety assurance and
commercial viability. Done poorly, it will create barriers that drive innovation offshore or
underground. WHM is confident that FDA, with thoughtful stakeholder engagement and
structural commitment to least-burdensome principles, will get this right.
We thank FDA for the opportunity to comment and welcome the opportunity to discuss any
aspect of this submission in a public meeting or pre-submission consultation context.
Respectfully submitted,
Walnut Hill Medical Healthcare Reimbursement & Commercialization Strategy Consulting Dallas,
TX August 18, 2026
END OF COMMENT