Hari Prakash Chanumolu
“A “competency-based approach” without a published evidence map is not a least-burdensome pathway; it is case-by-case review under a new name.”
What they argued
Q8 demands published risk-tier-to-evidence map; Q7 analogy justifies 'shape not quantity'; Q14 rejects median clinician; Q18 'concern'; PCCP 'around invariants', version pinning.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
Coded positions
Do not raise risk just because the user is a patient
Enforce limits on what the conversation can do
Require prospective studies for specified higher-risk uses
Check for shared blind spots in generated test data
Keep real evidence for claims synthetic data cannot establish
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Have clinicians review samples of outputs
Reassess after changes or safety signals
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Manage suitable changes through internal quality controls
Send specified changes back for FDA review
Specify the tests or controls a future change must pass
Define when a change needs further review
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Limit or test what the agent is allowed to do
Evaluate the full sequence of actions and its effects
Keep records that let investigators reconstruct actions
Across the five cross-cutting questions
High-consequence work: Advises
The comment as filed
See attached file(s)
Attachment
Public Comment
Docket No. FDA-2026-N-7874
Re: Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and
Request for Feedback
Submitted to: Division of Dockets Management, U.S. Food and Drug Administration, via Regulations.gov
Submitted by: Hari Prakash Chanumolu — submitted in an individual capacity
Date: August 18, 2026
Introduction and summary of position
I appreciate the opportunity to comment on CDRH’s discussion paper. I write in an individual capacity, from
a regulatory and quality-affairs perspective, based on work with software-as-a-medical-device submissions,
quality management systems, and change control. The views here are my own and do not represent any
employer or client.
The discussion paper is a serious and unusually candid document. Several of its questions — particularly
Questions 10, 13, 20, 23, and 24 — name problems that the field has been avoiding, and CDRH deserves
credit for putting them on the record rather than deferring them. My comments are offered in that spirit: I
support the general architecture the paper describes, and my criticisms are directed at the places where I
believe the framework, as currently sketched, would not survive contact with a real submission or a real
postmarket failure.
My principal recommendations:
1. Reconcile the two-axis risk framework with ISO 14971 rather than standing up a parallel
taxonomy. As drafted, the two axes omit the probability that an incorrect output escapes detection — the
variable that, for generative systems, drives most of the residual risk.
2. Publish an explicit mapping from risk tier to minimum evidence. A “competency-based approach”
without a published evidence map is not a least-burdensome pathway; it is case-by-case review under a
new name.
3. Do not permit synthetic data to establish performance rates. Synthetic data is a valid tool for finding
failures and covering rare events. It cannot supply a valid denominator, and when generated by a model of
the same family as the device under test it is structurally incapable of detecting shared blind spots.
4. Reject “median clinician in practice” as the primary performance standard. Prespecified clinical
acceptance criteria derived from the consequence of the decision are the better anchor; clinician panels
should serve as adjudication instruments, and should themselves be validated and reported like
instruments.
1 of 19
Docket No. FDA-2026-N-7874
5. Require independence for any supervisory agent offered as postmarket evidence. A monitoring agent
built on the same foundation model as the device it monitors has correlated failure modes and will be
blind to precisely the errors that matter.
6. Treat model version pinning as a design control expectation. For devices built on third-party
foundation models, the inability to pin a version-identified endpoint, with contractual change notice,
should be treated as a design deficiency rather than an accepted risk. This is the single largest unaddressed
gap in the paper.
7. Reframe the PCCP for generative systems around invariants rather than change lists. A PCCP that
prespecifies what must remain true — a locked evaluation instrument, locked thresholds, a locked
verification protocol — is workable where a PCCP that prespecifies what will change is not.
I. Assessment of risk (Section IV; Questions 1–6)
Question 1 — Dimensions of the risk framework
The two axes — device activity and the consequence of relying on an incorrect output — are necessary but
not sufficient. Together they describe severity and a proxy for autonomy. They do not describe the
probability that an incorrect output is acted upon, which is where generative systems differ most sharply
from the software CDRH has regulated to date.
This matters because ISO 14971, which manufacturers already apply and document, defines risk as the
combination of the probability of occurrence of harm and the severity of that harm. The paper’s framework
captures severity well and probability not at all. A device whose errors are obvious on inspection and a device
whose errors are fluent, confident, and internally consistent may occupy the same cell of the proposed grid
while presenting materially different risk. Fluency is precisely the property generative models optimize for,
and it is an anti-correlate of detectability.
I recommend CDRH add detectability — the likelihood that an incorrect output is recognized as incorrect
before it is relied upon — as an explicit dimension, and treat reversibility, downstream safeguards, time
pressure, and output traceability as determinants of detectability rather than as independent axes. This
preserves a two-dimensional grid that remains usable in practice while giving the omitted variable a defined
home.
I further recommend that CDRH state explicitly how this framework relates to a manufacturer’s existing ISO
14971 risk management file. If the answer is that the grid is a communication and triage device that draws on
the 14971 file rather than a separate analysis, that should be said. Absent such a statement, manufacturers will
reasonably assume they must maintain a second, FDA-specific risk taxonomy alongside the first, which
produces documentation burden without a corresponding safety gain.
2 of 19
Docket No. FDA-2026-N-7874
Question 2 — Directiveness as a continuum
I want to raise a practical objection to using directiveness as a determinant of risk classification,
notwithstanding that it is clearly relevant to risk.
Directiveness of a generative output is an emergent property of the output, not a design input the
manufacturer controls or can verify at design time. A manufacturer can constrain it — through system
prompts, output templates, refusal training, post-generation filtering — but cannot guarantee it. If regulatory
classification turns on whether a function is “non-directive,” sponsors will be asked to attest to something no
sponsor can honestly attest to, and the predictable result is either attestations that are not meaningful or an
evidentiary standard that cannot be met.
The more workable construction treats directiveness as a constrained and monitored property. A sponsor
would specify a directiveness ceiling for the intended use, describe the technical controls enforcing it, and —
critically — report a measured escape rate: the rate at which outputs exceeded the specified ceiling across an
adversarial and representative evaluation set. Classification would then be based on the specified ceiling,
conditioned on the escape rate meeting a prespecified bound, with escape rate becoming a postmarket
monitoring endpoint. This gives manufacturers something they can actually design toward and CDRH
something it can actually verify.
Question 3 — Patient-facing functions
I would caution against a categorical elevation of risk for patient-facing functions. The paper is right to flag
the tension, and right to warn against underestimating patient capability.
The variable that actually distinguishes the two cases is not the user’s domain knowledge in the abstract but
whether the user has a realistic path to verification at the moment of reliance. A clinician who lacks
specialty knowledge but can order a confirmatory test, consult a colleague, or check a reference is better
protected than a patient at home at 2 a.m. — but a patient reading a clearly sourced output with an explicit
instruction to call their clinician before acting is better protected than a clinician under time pressure
accepting a plausible output without checking.
I recommend the framework attend to the verification path — its availability, its latency, and whether the
output’s design supports or discourages it — rather than to the user’s category. Safeguards that meaningfully
change the analysis include traceable citation to primary sources with verifiable links, explicit uncertainty
communication, and a designed prompt to seek confirmation that scales with the consequence of the specific
output rather than appearing as static boilerplate.
Question 4 — Generalist versus specialist users
This distinction is analytically real but creates an enforcement problem CDRH should confront directly. A
manufacturer can label an intended user; it cannot control which clinician opens the application. Once a
device is deployed in an institution, use by clinicians outside the labeled specialty is foreseeable, and in some
settings it is the norm.
3 of 19
Docket No. FDA-2026-N-7874
If CDRH intends specialty scope to bear regulatory weight, the paper should distinguish between (a) labeled
intended user, (b) technical enforcement of user scope, and (c) foreseeable off-label user, and should state
which of the three drives the risk assessment. My recommendation is that risk be assessed against foreseeable
use, with credit given for demonstrated technical enforcement — role-based access controls tied to
institutional credentialing, for example — rather than for labeling alone. Labeling-only mitigations have a
poor track record in software and should not be treated as risk-reducing here.
Question 5 — Multi-turn conversational devices
The paper identifies the correct problem: intended use cannot be meaningfully characterized at the level of a
single turn when behavior is emergent across a conversation.
I recommend CDRH state that for multi-turn devices, the session, not the turn, is the unit of analysis for
both intended use and evaluation. Intended use should be defined at the conversational level, in terms of the
range of trajectories the device is designed to support and the boundaries it is designed to hold. Evaluation
should then be trajectory-based: multi-turn adversarial evaluation in which evaluators actively attempt to
walk the device from non-directive information toward action-directing output, with the escape rate measured
and bounded as described under Question 2.
Single-turn benchmarking of a multi-turn device measures something that does not correspond to how the
device is used, and I would encourage CDRH to say so plainly. Turn-level evaluation of conversational
devices should be treated as insufficient on its own regardless of how thorough it is.
Question 6 — Under- and over-escalation
The paper correctly observes that these two error directions may not be commensurable. My recommendation
is procedural: rather than asking manufacturers to reconcile them, require that the trade-off be prespecified
rather than discovered.
A sponsor should state, before evaluation, an explicit asymmetric acceptance criterion — for example, a
minimum sensitivity for the escalation-warranted condition together with a maximum acceptable overescalation rate — accompanied by a clinical justification for the chosen ratio that references the specific
deployment context and the capacity of the receiving system. The justification, not just the numbers, should
be part of the submission and should be assessed. Post hoc reporting of both rates without a prespecified
trade-off allows the sponsor to characterize whichever result is more favorable as the primary endpoint, and
CDRH should foreclose that.
Over-escalation deserves more attention than it typically receives. Its harms — alert fatigue, resource
diversion, downstream harm to patients who did need the capacity consumed — are diffuse, delayed, and
rarely attributed to the device. They are therefore systematically under-detected by exactly the postmarket
mechanisms Section VI describes.
4 of 19
Docket No. FDA-2026-N-7874
II. Premarket evaluation and the competency-based approach (Section V; Questions 7–
17)
Question 7 — Is the competency-based approach appropriate?
Broadly, yes, and the paper’s framing is a genuine contribution. But the analogy carries an assumption that
should be made explicit before it is relied upon, because I believe the analogy is doing more argumentative
work than it can support.
Physician credentialing is evidentiarily permissive — no exhaustive testing of every scenario — because it is
embedded in a system of continuous individual accountability: ongoing licensure with revocation authority,
institutional privileging and peer review, malpractice liability, professional norms, and a career-long incentive
to notice and correct one’s own errors. The credentialing exam is not the safety mechanism. It is the entry
gate to a system whose safety mechanisms operate continuously thereafter.
A device inherits none of this. There is no analogue to license revocation for a deployed model, no peer
review of its individual decisions, no professional identity that responds to being wrong. The manufacturer’s
quality system is the closest analogue, and it is a genuine one, but it is a different mechanism with different
failure modes.
There is a second disanalogy worth stating. Physicians generalize from training in ways whose failure modes
are broadly legible to other physicians — we have vocabulary for the kinds of errors a tired resident makes.
Generative model failures are not legible in the same way. They can be sharp, input-specific, and invisible to
any amount of aggregate performance measurement.
My recommendation is that CDRH retain the competency framing as an organizing structure — it maps
cleanly onto benchmarking, confirmation, and monitoring — while explicitly declining to import its
evidentiary permissiveness. The paper should state that the licensure analogy justifies the shape of the
evidence, not the quantity, and that the reduced-testing feature of physician credentialing is a consequence of
accountability structures the device context does not reproduce.
Question 8 — Relating the risk framework to evidence requirements
This is, in my view, the most consequential question in the paper for whether the framework succeeds in
practice.
A competency-based approach without a published risk-tier-to-evidence mapping is not a least-burdensome
pathway. It is discretionary, case-by-case review with a new vocabulary. Sponsors cannot plan submissions
against it, cannot budget against it, and cannot know in advance whether their evidence is adequate — which
in practice means either over-generating evidence defensively or under-generating it and absorbing review
cycles. Both outcomes are costly, and the second is worse for patients because it delays devices that would
have qualified.
I recommend CDRH publish a mapping table specifying, for each cell or band of the risk grid: the minimum
benchmarking elements required; the minimum acceptable clinical confirmation tier from the Section V.C
5 of 19
Docket No. FDA-2026-N-7874
ladder; whether third-party involvement is expected; and the minimum postmarket monitoring configuration.
The mapping need not be rigid — a documented, justified deviation pathway should exist — but the default
must be published and stable. Predictability is the mechanism by which a risk-based framework actually
reduces burden.
Question 9 — Adequacy of the benchmarking structure
The categories described appear substantially complete. I note three gaps.
Adversarial and hostile input. Robustness and reliability, as described, appear oriented toward natural
variation in input — phrasing, formatting, incomplete data. Generative devices additionally face adversarial
input: prompt injection through retrieved documents or patient-supplied text, instruction override, and
jailbreak techniques that defeat scope constraints. This is a security property, not a robustness property, and it
should be a distinct benchmarking element cross-referenced to CDRH’s premarket cybersecurity
expectations. A device whose scope boundaries hold under natural input and fail under a prompt embedded in
an uploaded document has not demonstrated boundary adherence.
Provenance and citation fidelity. Where a device cites sources, the correctness of the citation is a distinct
failure mode from the correctness of the claim. A fabricated but plausible citation actively defeats the user’s
verification path — it converts a safeguard into a hazard, because the user who checks is reassured by the
presence of a reference. Citation fidelity should be measured separately and should not be assumed to track
content accuracy.
Behavior at the edge of scope. Scope maintenance measures whether the device stays in scope. Equally
important is what it does at the boundary: whether performance degrades gracefully with visible uncertainty,
or falls off sharply while confidence remains high. The latter is far more dangerous and is invisible to
aggregate accuracy metrics.
On external standards: I am aware of ongoing work in AI evaluation and risk management standards that may
be relevant here, but I do not have current verified information on which specific standards have reached a
maturity appropriate for regulatory recognition, and I would not want to cite one incorrectly. I recommend
CDRH survey this landscape directly and state which standards it considers suitable for leveraging.
Question 10 — Benchmark validity, contamination, and sponsor-developed benchmarks
This question deserves the most attention of any in the paper, because benchmarking is load-bearing for the
entire premarket structure and the assets currently available cannot bear that load.
Contamination is not testable without a training-data cutoff attestation. A sponsor cannot demonstrate
that a public benchmark was absent from a third-party model’s training corpus without information the
sponsor does not possess. I recommend CDRH state that where a device relies on a third-party model,
benchmarking evidence based on publicly available assets is not interpretable absent a dated training-data
cutoff attestation from the model provider, and that temporally held-out evaluation data — constructed from
material post-dating the attested cutoff — is the preferred mitigation. This also creates the market pull
discussed under Question 25.
6 of 19
Docket No. FDA-2026-N-7874
Construct validity must be argued, not assumed. I recommend CDRH require a sponsor to state, for each
benchmark used, the specific real-world behavior it is intended to predict, the reasoning connecting them, and
the known limitations of that connection. This is a familiar requirement in a different vocabulary: it is
analytical validation, and it should be documented to the same standard.
Sponsor-developed benchmarks should be permitted but structurally separated. They are often
unavoidable — for a novel intended use, no external benchmark exists. The controlling concern is not that the
sponsor built the benchmark but that the sponsor may have optimized against it. I recommend three
safeguards: (a) the benchmark and its acceptance thresholds are locked and submitted before evaluation, not
after; (b) the sponsor attests, subject to audit, that the gating set was not used for model selection, prompt
engineering, fine-tuning, or any iteration on the device — the separation between development and gating
sets must be demonstrable from development records, not merely asserted; (c) a held-out portion is escrowed
with FDA or a recognized third party and never returned to the sponsor.
Public benchmarks should be characterized as screening, not evidence. Given contamination and
saturation, strong performance on a public benchmark establishes little; weak performance establishes a great
deal. I recommend CDRH treat public benchmarks as necessary-but-not-sufficient screens and state explicitly
that they do not constitute evidence of effectiveness.
A rotating non-public evaluation resource is worth serious consideration. The contamination problem is
structurally unsolvable by sponsors acting individually, since any benchmark that becomes valuable becomes
public and then becomes training data. A periodically refreshed evaluation set, maintained by FDA or a
recognized body and never published, is one of the few mechanisms that addresses the root cause. I recognize
this carries substantial resource and governance implications, and I raise it as a direction worth evaluating
rather than as a costed proposal.
Question 11 — Selecting a clinical confirmation approach
I support the graduated ladder. Two recommendations.
First, the selection should be driven by the risk mapping recommended under Question 8, not left to
sponsor justification. Sponsor discretion in selecting one’s own evidentiary bar, even with a justification
requirement, produces predictable downward pressure and inconsistent review.
Second, shadow deployment deserves to be the presumptive default at intermediate risk. It is the only
rung that produces genuine real-world input distributions at zero patient exposure, and the divergence
between benchmark input distributions and real clinical input distributions is, in my experience, where
software devices most often disappoint. Retrospective evaluation on curated real inputs does not substitute,
because curation removes exactly the malformed, incomplete, and atypical inputs that drive real-world
failure.
On input distributions: I recommend CDRH require prospective characterization of the anticipated real-world
input distribution, and that confirmation evidence be labeled with the population and setting in which it was
obtained. A device confirmed at three academic centers has not been confirmed for a rural critical-access
7 of 19
Docket No. FDA-2026-N-7874
hospital with different documentation practices, and the labeling should make the scope of the evidence
visible to the adopting institution.
Question 12 — Statistically meaningful performance measurement
There is a measurement problem here that classical device statistics do not address and that the paper does
not raise: generative devices are non-deterministic at fixed input. The same input can produce different
outputs across runs, and in some cases outputs that differ in clinical direction. Conventional performance
estimation assumes a fixed input-output mapping and therefore attributes all observed variance to case mix.
I recommend CDRH require:
• Repeated-sampling variance reporting. Performance must be characterized across repeated runs at
fixed input under the deployed sampling configuration, with within-input variance reported separately
from between-input variance.
• Acceptance criteria stated on a lower confidence bound. For a non-deterministic device, a point
estimate meeting a threshold does not establish that the device meets the threshold. The bound, not the
estimate, should be the criterion.
• Sampling configuration treated as a specified device parameter subject to change control.
Temperature, top-p, and equivalent settings materially change the performance distribution, and a device
evaluated at one configuration and deployed at another has not been evaluated.
On combining benchmarking and confirmation evidence into a single estimate: I recommend against
permitting it. The two arise from different sampling frames and support different inferences —
benchmarking establishes capability under constructed conditions, confirmation establishes performance
under real ones. Pooling them produces an estimate that describes no actual population, and it allows large
volumes of cheap benchmark data to dominate small volumes of expensive confirmation data, diluting
precisely the evidence that matters most. They should be reported separately, with the confirmation estimate
governing the acceptance decision and the benchmarking estimate serving as supporting evidence and as the
locked postmarket baseline.
Question 13 — Synthetic data
The paper’s framing of this question — the risk that synthetic data generated by models of the same class as
the device reproduces the very gaps evaluation should detect — is exactly right, and I want to press on the
implication.
The concern is not bias in the usual sense; it is correlated failure. A generator sharing training data,
architecture, or lineage with the device under test shares its blind spots. Where the device misunderstands a
rare presentation, a same-family generator will produce synthetic cases embodying the same
misunderstanding. The evaluation then returns a clean result on a population that does not exist, and the
failure is invisible because the evidence looks complete. This is worse than having no synthetic data, because
it manufactures unwarranted confidence.
8 of 19
Docket No. FDA-2026-N-7874
I recommend CDRH draw a bright line by purpose rather than by domain:
• Synthetic data is appropriate for finding failures. Stress testing, adversarial probing, coverage of rare
and dangerous scenarios that cannot be ethically or practically collected. Here a failure discovered is
informative regardless of whether the input was real, and the correlated-blind-spot problem biases toward
under-detection — a conservative direction.
• Synthetic data is not appropriate for establishing performance rates. Any rate requires a denominator
representing a real population. Synthetic data cannot supply one, because the generator’s distribution is
unvalidated against the deployment distribution and unvalidatable except by reference to the real data one
is trying to avoid collecting.
I further recommend that same-family generation be prohibited for subgroup performance claims
specifically. Subgroup performance is where correlated blind spots do the most damage — underrepresented
subgroups are underrepresented in the generator’s training data for the same reasons they are
underrepresented in the device’s — and it is where a false negative has the clearest equity consequence.
Where synthetic data is used at all, the generator’s identity, provenance, and training lineage should be
disclosed, and its relationship to the device’s underlying model stated explicitly.
I am not aware of a validated method for establishing that a synthetic clinical population is distributionally
adequate to a real one for regulatory purposes, and would be interested to see any that commenters can point
to. I would rather CDRH proceed as though none exists than adopt a permissive posture on the assumption
that one will emerge.
Question 14 — Comparators and acceptance criteria
I recommend against “the median clinician in practice” as a primary performance standard, for several
reasons.
It is unmeasured — there is no reliable characterization of median clinician performance for most tasks a
GenAI device would perform, so the comparator would in practice be a convenience sample presented as a
population parameter. It is a moving target, and one that a widely adopted device would itself move. Most
importantly, it anchors the acceptable bar to observed practice variation rather than to patient benefit: where
median practice is poor, the standard licenses a device that is also poor, and the patient in front of that device
receives no protection from the fact that the alternative was equally bad.
The better construction, in my view, has two parts:
1. Prespecified clinical acceptance criteria derived from the consequence of the decision. What
performance is required for this device, in this use, to be safe and effective, given what happens when it is
wrong? This is a clinical and normative judgment, it should be made in advance and defended, and it is
the judgment the acceptance decision actually turns on.
2. A clinician panel as an adjudication instrument, not as the bar. Panels are essential for open-ended
outputs, where automated scoring is inadequate. But a panel is a measurement instrument, and it should
be validated and reported as one: prespecified composition and qualifications, prespecified size with
9 of 19
Docket No. FDA-2026-N-7874
justification, a written adjudication protocol, blinding to output source where feasible, measured and
reported inter-rater reliability, and a prespecified procedure for resolving disagreement. A panel with
unreported inter-rater reliability is an uncharacterized instrument, and evidence generated by an
uncharacterized instrument should not gate a marketing decision. This is standard practice in reader
studies and should carry over without modification.
On generalist versus specialist: the comparator should match the labeled intended user, subject to my
comment under Question 4 regarding foreseeable use. A device labeled for primary care benchmarked against
subspecialists is being measured against a standard its users do not represent — which may overstate or
understate the device’s contribution depending on direction, and in either case measures the wrong thing.
On human-AI team performance: where the labeling requires human review, team performance should be
the primary basis of evaluation, and standalone device performance should be supporting evidence only.
Evaluating the device in isolation when it will never be used in isolation measures a configuration that does
not exist. I would add that team evaluation must be designed to detect automation bias — the tendency of
reviewers to defer to a confident output — which is a principal mechanism by which human-in-the-loop
safeguards fail. A team evaluation that does not include cases where the device is confidently wrong will
systematically overstate the safeguard’s value.
Question 15 — Counterfactual comparators
I support permitting comparison to the care that would occur absent the device, and I think the paper is right
that this is often the clinically meaningful question. For access-expanding devices, the true alternative is
frequently no assessment at all, delayed assessment, or assessment by a less qualified party, and holding such
a device to a specialist standard can foreclose a real benefit.
Two safeguards should accompany it. First, the counterfactual must be empirically characterized for the
specific deployment setting, not assumed — the claim that “nothing would otherwise happen” is an
evidentiary claim about a care pathway and should be supported. Second, and more importantly, the
comparator must travel with the labeling. A device authorized against a no-intervention counterfactual in a
resource-limited setting has not been shown safe and effective as a substitute for specialist review in a
resource-rich one, and the authorization should say so in terms an adopting institution can act on. Without
this, counterfactual comparators become a low-bar entry route followed by scope creep that no one has
evaluated.
Question 16 — Independent third parties
Third-party involvement is well suited to benchmark administration, escrow of held-out evaluation sets, and
clinical adjudication — activities where independence from the sponsor is the substantive point.
The design risks are real and familiar from other accreditation regimes. I recommend:
• Published recognition criteria for qualifying bodies, so recognition is not discretionary and entry is
contestable.
10 of 19
Docket No. FDA-2026-N-7874
• Strict separation of advisory and assessment roles. An entity that advised a sponsor on device
development, benchmark design, or submission strategy must not assess that device. This is elementary
quality-system independence and should be stated as a disqualification, not a disclosure.
• Public disclosure of financial relationships between assessing bodies and sponsors, including aggregate
revenue concentration — a body deriving most of its revenue from a small number of sponsors is not
independent in any meaningful sense regardless of per-engagement firewalls.
• A preserved first-party pathway. Third-party assessment should be an option that reduces review
friction, never a mandatory gate. Making it mandatory converts recognized bodies into rent-collecting
chokepoints and disproportionately burdens smaller developers, which is a competition harm and, over
time, a safety harm through reduced diversity of approaches.
I would also flag that third-party assessment capacity for generative medical devices does not currently exist
at scale in any form I am aware of. If CDRH intends third parties to play a significant role, the paper should
acknowledge the capacity-building timeline, because a framework that depends on institutions that do not yet
exist will not be operable on the timeline the technology is moving.
Question 17 — Other model architectures
I would encourage caution in extending the framework by analogy. The competency-based approach is
workable for language-centric devices in part because a meaningful, if flawed, ecosystem of language
benchmarks exists. For multimodal vision-language models and generative or predictive world models, the
benchmarking assets are substantially thinner, and I do not have verified information on what is currently
available for clinical use cases in these modalities.
The structural implication is that where benchmarking assets are immature, clinical confirmation must
carry proportionally more weight. I recommend CDRH state this scaling principle explicitly rather than
allowing a uniform framework to imply that a thin benchmarking record is as informative as a rich one.
Otherwise the framework’s evidentiary flexibility will be claimed most aggressively exactly where the
underlying evidence base is weakest.
III. Postmarket monitoring (Section VI; Questions 18–24)
Question 18 — Trading premarket evidence for postmarket monitoring
This is an appealing proposition and it may be the right direction, but I want to register a concern grounded in
how postmarket commitments have historically performed.
The general pattern across regulated products is that postmarket obligations are less reliably fulfilled than
premarket ones — completion is slower, enforcement is resource-intensive, and passive surveillance systems
depend on reporting behavior that is uneven. I do not have current verified figures on postmarket study
completion rates or software-related adverse event reporting rates, and I would encourage CDRH to publish
its own data on this before relying on postmarket monitoring as a substitute for premarket evidence. If the
11 of 19
Docket No. FDA-2026-N-7874
data show that postmarket obligations for software devices are reliably met, that materially strengthens the
case; if they do not, the paper’s proposal transfers risk to patients in exchange for an obligation that may not
be discharged.
Conditions I would consider necessary:
• The monitoring program must be an enforceable condition of authorization, with prespecified
triggers and prespecified consequences — including suspension of marketing — that attach automatically
rather than through a discretionary enforcement decision. A commitment is not a control.
• The sponsor must demonstrate technical capability to observe the endpoint. Many GenAI devices
have no linkage between an output and any recorded clinical outcome. Where the sponsor cannot show a
data pathway from output to observable consequence, the monitoring program is aspirational and cannot
justify reduced premarket evidence. This should be an explicit gating question.
• Detection latency must be shorter than harm accumulation. Where a degradation would produce harm
faster than the monitoring cadence could detect it, the trade is unavailable regardless of program quality.
• Irreversibility forecloses the trade. Where the harm from an incorrect output cannot be undone,
postmarket detection is not a mitigation — it is an audit of damage already done.
I recommend CDRH state affirmatively that reduced premarket evidence is not available for devices in the
high-consequence, high-activity region of the risk grid, so that the flexibility is bounded on its face.
Question 19 — Approaches and cadence
The three approaches described are sensible and complementary. My substantive recommendation concerns
cadence: it should be event-driven with a calendar floor, not calendar-driven.
Calendar-based reassessment is poorly matched to a technology whose performance changes discontinuously
at the moment of a change rather than continuously with time. I recommend triggering events include, at
minimum: any change to the underlying model version or weights, whether initiated by the manufacturer or
the model provider; any change to system prompts, instructions, or output templates; any change to a retrieval
corpus or knowledge source; any measured shift in the input distribution beyond prespecified bounds; any
change in the deployed sampling configuration; and any accumulation of complaints or adjudicated errors
exceeding a prespecified threshold. A calendar floor — annual, or more frequent for higher-risk devices —
should catch slow drift that trips no discrete trigger.
From a quality-system perspective these should be defined as change-control triggers within the
manufacturer’s QMS, so that the postmarket monitoring obligation is integrated with existing design-change
procedures rather than maintained as a separate process. Parallel processes diverge.
Question 20 — Machine-based supervisory agents
Automated supervision is probably unavoidable at the volumes involved — human review of every output is
not feasible for a high-throughput device — so the question is how to make it trustworthy rather than whether
to permit it.
12 of 19
Docket No. FDA-2026-N-7874
The dominant risk is correlated failure. A supervisory agent built on the same foundation model as the
device it monitors shares that model’s blind spots. It will approve exactly the outputs that are wrong for
reasons the model family does not represent, and its agreement will be misread as independent confirmation.
This is not a hypothetical: it is the direct consequence of shared training data and architecture, and it means a
same-model supervisor’s clean report carries close to zero information about the failures that matter most.
I recommend:
• Architectural independence as a condition of evidentiary use. A supervisory agent offered as
postmarket evidence should be built on a model of a different family and provider than the supervised
device. Where full independence is impractical, the sponsor should characterize and report the correlation
in failure modes between supervisor and device on a held-out error set, and the monitoring program’s
credit should scale with the demonstrated independence.
• The supervisor is itself a measuring instrument and must be validated as one. Report its sensitivity
for the specific harms of interest against human adjudication on a sampled basis — not aggregate
agreement, which is dominated by the easy majority of correct outputs and can be high while sensitivity
for serious errors is near zero. Sensitivity must be re-measured over time, since the supervisor’s own
performance drifts.
• The supervisor triages; it does not conclude. Its appropriate role is prioritizing outputs for human
review and enabling broader sampling than humans could achieve unaided. The adjudicated endpoint
should remain human. A monitoring program whose terminal judgment is machine-made has no
independent check anywhere in the loop.
• The supervisor is part of the device system for change-control purposes. A change to the supervisory
agent changes the device’s monitoring properties and should trigger reassessment on the same terms as a
change to the device.
Question 21 — Ecosystem roles and manufacturer accountability
The accountability principle can be stated simply, and I recommend CDRH state it: information may come
from anywhere; the duty to investigate and act remains with the manufacturer. Clinicians, institutions,
and societies can be valuable sources of signal, and health systems in particular hold the outcome data
manufacturers lack. But a manufacturer that has arranged to receive signal from third parties has not thereby
delegated any portion of its complaint-handling, investigation, or reporting obligations, and the paper should
say so directly to prevent the shared-ecosystem framing from being read as shared responsibility.
I recommend building on the existing complaint-handling and MDR architecture rather than creating a
parallel structure. That said, I want to flag a gap that I believe requires attention before postmarket
monitoring can function at all for these devices:
It is currently unclear what constitutes a reportable event for a generative output. If a device produces a
clinically incorrect recommendation and the clinician recognizes and disregards it, has a device malfunction
occurred? Under a conventional reading, the device did not perform as intended, and a malfunction that could
cause or contribute to serious injury if it recurred is reportable — which would imply that hallucinations
13 of 19
Docket No. FDA-2026-N-7874
caught by users are reportable events. I do not believe manufacturers are currently operating on that reading,
and I do not believe the volume would be manageable if they did. But the alternative reading — that only
outputs resulting in actual harm are reportable — discards precisely the near-miss data that would make
postmarket monitoring effective, and near-miss data is the most valuable signal any surveillance system
collects.
This ambiguity is consequential and unresolved, and every element of Section VI depends on how it is
settled. I recommend CDRH address it directly, and would suggest that the workable answer likely involves a
distinct, aggregate reporting channel for caught errors — reported as rates rather than as individual MDRs —
so that the signal is preserved without generating unmanageable individual-event volume.
Question 22 — Scaling re-evaluation to modifications
I support using the premarket competency assessment as a locked baseline. The organizing principle I would
propose for categorizing changes is this: the question is not how large the change is, but whether the
locked evaluation instrument remains valid for the changed device. A small change that invalidates the
benchmark suite requires more scrutiny than a large change that does not.
Applying that principle, a tiered structure might look like:
Tier 1 — no re-benchmarking. Changes that cannot affect device outputs: user interface presentation,
logging, analytics instrumentation, infrastructure changes with no path to the output. Documented under
normal change control.
Tier 2 — full re-benchmarking against the locked suite, managed within the QMS, no premarket
submission. Model version updates within the same family and provider; prompt or template modifications
within a validated structure; retrieval corpus refreshes drawing on a validated source list; sampling
configuration changes within a validated range. The condition is that the device passes the complete locked
benchmark suite at a prespecified non-inferiority margin, with results documented and available on
inspection. Failure at any element converts the change to Tier 3.
Tier 3 — premarket submission. Change of foundation model family or provider; expansion of intended
use, intended user, or clinical scope; addition of a new input or output modality; any change to agentic tool
access or action scope; and any change for which the locked benchmark suite is no longer a valid instrument
— for example, a change that introduces capabilities the suite does not probe.
The essential safeguard is that the benchmark suite itself cannot be modified within Tier 2. If the sponsor
can revise the instrument and the device together, non-inferiority against the revised instrument is
meaningless. Modification of the locked suite should always be a Tier 3 event.
Question 23 — PCCPs when modifications cannot be prespecified
This is the right question and I think it admits a clean answer, which is the most useful thing I can offer in this
comment.
14 of 19
Docket No. FDA-2026-N-7874
The current PCCP construct prespecifies what will change — the modification, the methods to implement
and validate it, the impact assessment. For generative devices this fails, not because sponsors are unwilling
but because the change space is genuinely open: a sponsor cannot enumerate in advance the model versions a
provider will release or the prompt refinements that experience will suggest.
The adaptation is to invert what is locked. An invariant-based PCCP prespecifies not the change but what
must remain true after any change:
• The evaluation instrument — the complete benchmark suite, escrowed and immutable.
• The acceptance thresholds, including subgroup-level thresholds and the non-inferiority margin.
• The verification protocol — what is run, in what order, by whom, with what independence, before any
change reaches users.
• The rollback criteria and mechanism, including maximum time-to-rollback and the technical
demonstration that rollback is achievable.
• The change categories in scope, defined by the Tier 2 boundary above rather than by enumeration of
specific changes.
• The documentation and notification obligations attaching to each change.
Under this structure a sponsor may make any change falling within the scoped categories, provided the
changed device passes the locked evaluation at the locked thresholds under the locked protocol, with
everything documented and inspectable. The sponsor gains the operational flexibility the technology requires.
CDRH retains a fixed, auditable gate that does not move.
The construct depends entirely on the integrity of the locked instrument. If the sponsor can modify the
benchmark suite while modifying the device, the whole thing is circular and provides no assurance at all.
Escrow of the suite, and Tier 3 treatment of any change to it, are not optional features of this proposal — they
are the proposal.
Question 24 — Third-party model changes initiated by the model developer
In my assessment this is the most significant unaddressed gap in the current regulatory picture, and I am glad
the paper names it.
My understanding of prevailing commercial practice — which I would encourage CDRH to verify directly
with model providers, as it varies by provider and is changing quickly — is that hosted model endpoints are
frequently updated without individualized customer notice; that version pinning is offered inconsistently and
often with limited-duration guarantees; and that pinned versions are deprecated on provider-determined
timelines that may be shorter than a device lifecycle. If that characterization is broadly accurate, then a device
manufacturer relying on a hosted third-party endpoint may have its device’s behavior changed without its
knowledge and without any mechanism to detect the change, which is not a condition any quality system can
accommodate.
I recommend:
15 of 19
Docket No. FDA-2026-N-7874
Treat version pinning as a design control expectation. CDRH should state that a GenAI-enabled device
relying on a third-party model is expected to execute against a pinned, version-identified model endpoint, and
that the inability to pin — or reliance on a provider that does not offer pinning with change notice — is a
design deficiency to be addressed, not a residual risk to be accepted. This is the single highest-leverage
statement CDRH could make in this area, and it would immediately reshape provider offerings in the medical
device segment.
Require contractual change notice with a defined minimum period, sufficient to complete the Tier 2 rebenchmarking described above before a change takes effect, together with a minimum deprecation window
for pinned versions.
Require continuous canary monitoring, because contracts cannot be verified in real time. A small, fixed
probe suite — a set of inputs with known expected outputs — executed against the live endpoint at high
frequency provides direct detection of silent upstream change. This is inexpensive, technically
straightforward, and the only mechanism I am aware of that verifies rather than assumes endpoint stability. I
recommend CDRH describe it as an expected element of postmarket monitoring for any device on a thirdparty hosted model.
Treat provider-forced migration as a Tier 3 change. When a provider deprecates a pinned version and the
manufacturer must migrate, the resulting device is running on a different model. That it was involuntary does
not change what it is, and the paper should foreclose the argument that forced migrations warrant lighter
treatment. Manufacturers should be expected to plan for this contingency — including maintaining a
validated fallback — as part of design planning rather than handling it as an emergency.
I would also note that this is where a Foundation Model MAF (Question 25) would deliver the most value: a
standing, current record of a model’s versioning, deprecation, and change-notification practices would let
CDRH and manufacturers assess dependency risk directly rather than by inference.
IV. Other topics (Section VII; Questions 25–26)
Question 25 — Voluntary Foundation Model Master Files
The paper’s candor about the incentive problem is warranted. Foundation model developers serve markets in
which medical devices are a small fraction of revenue; disclosure of safety-relevant limitations carries
competitive and litigation exposure; and a voluntary program asking them to document weaknesses for a
regulator overseeing someone else’s product is unlikely to attract broad participation on its own merits.
I nonetheless recommend CDRH create it, because the mechanism that makes it work is not the model
developer’s incentive — it is the device manufacturer’s.
Pair the voluntary MAF with a manufacturer-side default: where an MAF is absent, the device
manufacturer bears the full characterization burden for the underlying model. Training data provenance
to the extent determinable, contamination assessment, known failure modes, version and deprecation
16 of 19
Docket No. FDA-2026-N-7874
practices, subgroup performance characterization — all of it, generated independently at the manufacturer’s
cost, or the submission is incomplete.
That default creates a procurement preference. Device manufacturers will strongly prefer model providers
that maintain a current MAF, because the alternative is expensive and often technically impossible. Providers
will then face a commercial reason to participate that has nothing to do with regulatory goodwill. This is how
a voluntary program acquires teeth without new authority: it does not compel the model developer, it makes
the device manufacturer want to compel them, and the device manufacturer is the party with the contract.
Suggested MAF content: model identity and version lineage; training data cutoff date (essential for the
contamination analysis under Question 10); versioning, change notification, and deprecation policy (essential
for Question 24); known limitations and documented failure modes; safety training and refusal behavior;
evaluation results the developer has generated; and any medical-domain-specific training or fine-tuning. A
currency requirement — the MAF must be updated within a defined period of any model change — is what
makes it useful rather than a snapshot.
I would add one caution: an MAF should never be treated as establishing the safety or effectiveness of any
device built on the model. The paper’s statement that the unit of evaluation is the final user-facing device as
configured for deployment is correct and important, and the MAF should be framed as context for that
evaluation, not as a partial substitute for it.
Question 26 — Agentic AI systems
Agentic systems present at least four considerations that do not arise, or arise much less acutely, for nonagentic generative devices.
Trajectory-level acceptance criteria. Per-step performance is a misleading summary of end-to-end
reliability. A device that completes each step correctly 95% of the time completes a ten-step task correctly
about 60% of the time, and that arithmetic is unforgiving as sequences lengthen. Acceptance criteria must be
stated at the level of the completed task, and evaluation must measure task completion rather than step
accuracy. I recommend CDRH state this explicitly, because step-level metrics are far easier to generate and
will otherwise be offered.
Action inventory with irreversibility classification. I recommend sponsors submit an enumerated inventory
of every action the system can take — every tool, every write operation, every external system it can affect
— each classified by reversibility and by the consequence of erroneous execution. Irreversible actions should
require explicit confirmation by a qualified user, and the inventory should be treated as a locked specification
whose expansion is a Tier 3 change under the framework in Question 22. Uncontrolled expansion of an
agent’s action scope is the most likely path from a low-risk deployment to a high-risk one without any
regulatory event marking the transition.
Least-privilege tool access. Agentic systems should hold the narrowest permissions sufficient for the
intended use, scoped per deployment. This is standard security practice and maps directly onto CDRH’s
existing cybersecurity expectations; I recommend the paper cross-reference them rather than developing a
parallel treatment.
17 of 19
Docket No. FDA-2026-N-7874
Audit logging sufficient for post-hoc reconstruction. When an agentic system produces a harmful outcome,
the investigation must be able to reconstruct the full trajectory — inputs, intermediate reasoning artifacts
where available, tool calls, returned values, actions taken, and the model version and configuration in effect.
Without this, root cause analysis is not possible, corrective action is guesswork, and the complaint-handling
obligations under Question 21 cannot be discharged. I recommend audit logging adequacy be an explicit
premarket review element for agentic devices rather than a postmarket assumption.
I would also note that the compounding-error arithmetic above interacts badly with the premarket-uncertainty
trade proposed in Question 18. Agentic devices are where end-to-end performance is hardest to establish
premarket, and simultaneously where errors propagate fastest and are least reversible postmarket. I
recommend CDRH treat agentic devices as presumptively outside the reduced-premarket-evidence pathway.
V. Cross-cutting recommendations
1. Publish the risk-tier-to-evidence mapping. Predictability is the mechanism by which risk-based
frameworks reduce burden. Without it, “competency-based” and “least burdensome” will not describe the
same thing.
2. Integrate with existing quality system and risk management infrastructure. Manufacturers maintain
ISO 14971 risk files, design controls, and change control procedures under the QMSR. The framework
should draw on these rather than establishing parallel artifacts, and the paper should state how the pieces
relate.
3. Resolve MDR reportability for generative outputs. Every element of the postmarket program depends
on it, and it is currently ambiguous in a way that produces both under-reporting and unmanageable
exposure depending on how a manufacturer reads it.
4. State a design control expectation for model version pinning. This is achievable now, addresses the
largest structural gap in the framework, and would immediately improve provider practices in the medical
device segment.
5. Identify which elements require new legal authority. The paper expressly defers this question. I
understand why, but I would encourage CDRH to return to it, because several proposals — enforceable
postmarket conditions with automatic consequences, third-party recognition regimes, obligations touching
foundation model developers who are not device manufacturers — appear to sit at or beyond the edge of
existing authorities. Stakeholders assessing which parts of this framework to build toward would benefit
from knowing which parts CDRH can implement and which would require Congress.
6. Preserve the option to say no. The paper’s framing is oriented toward enabling market entry,
appropriately so given its purpose. But a competency-based framework should also be capable of
concluding that a given intended use is not currently supportable by available evidence. I would
encourage CDRH to state that the framework contemplates that outcome, so that it is understood as an
evaluation framework rather than an authorization pathway.
18 of 19
Docket No. FDA-2026-N-7874
Closing
I appreciate CDRH’s willingness to publish preliminary thinking and invite criticism of it before positions
harden. The competency-based framing is a real contribution, and the paper’s questions are, in several places,
sharper than the field’s current answers.
My central concern is that the framework’s flexibility currently outruns the infrastructure needed to make
flexibility safe: benchmarks that cannot yet be trusted, postmarket monitoring whose reportability rules are
unsettled, third-party model dependencies that manufacturers cannot presently control, and third-party
assessment capacity that does not yet exist. Each of these is addressable, and several are addressable now. But
the flexibility and the infrastructure should arrive in that order — infrastructure first — rather than the
reverse.
I would be glad to provide further detail on any point, and I thank the Agency for its consideration.
Respectfully submitted,
Hari Prakash Chanumolu
August 18, 2026
Submitted in an individual capacity. The views expressed are my own and do not represent those of any
employer, client, or organization.
19 of 19