Sehouenou Alberic Candide Ahouehome
What they argued
Supports least-burdensome evidence gradient; Q18 rebalancing ill-suited for autonomous severe-consequence functions; requires predictive-validity and sequestered sets; 'change-envelope' PCCP prespecifying verification protocol.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ26 · Agentic devices
Coded positions
Consider how personalized the answer is
Do not raise risk just because the user is a patient
Require specialist review or escalation when needed
Test with the intended clinician group
Define when the AI must escalate or defer
Protect test sets from exposure or contamination
Across the five cross-cutting questions
High-consequence work: Advises
The comment as filed
The discussion paper reflects a careful, risk-proportionate, least-burdensome orientation that I support. My principal recommendations are: (i) keep the risk framework two-dimensional and operationalize the additional dimensions, including error detectability by the user, as documented modifiers on the consequences axis; (ii) publish an illustrative mapping from risk position to expected evidence, and an anchored directiveness rubric applied to the empirical distribution of device outputs; (iii) require predictive-validity evidence for gating benchmarks, with contamination testing and independent sequestered test sets; (iv) encourage shadow deployment and require formal transportability assessment for RWD and OUS evidence; (v) anchor subgroup performance claims in a minimum core of real data; (vi) pair behavioral postmarket monitoring with outcome-linked surveillance in real-world data, with event-driven re-benchmarking cadence; (vii) adapt PCCPs toward verification-protocol (“change-envelope”) prespecification for changes that cannot be enumerated in advance; and (viii) pursue international alignment of these concepts, including MAF content, with partner regulators and through IMDRF.
Thank you for considering these comments. I would welcome any opportunity to provide further input.
Respectfully submitted.
Attachment
August 27, 2026
Sehouenou Alberic Candide Ahouehome, M.Sc. Epidemiology, MBA business Analytic
Program Evaluator, Québec City, Québec, Canada
candide.ahouehome.1@ulaval.ca
Dockets Management Staff (HFA-305)
Food and Drug Administration
5630 Fishers Lane, Rm. 1061, Rockville, MD 20852
Submitted electronically via https://www.regulations.gov
Re: Docket No. FDA-2026-N-7874: “Considerations for the Regulation of Generative AI-Enabled Medical
Devices: Discussion Paper and Request for Feedback” (issued August 18, 2026; comments due October
19, 2026)
Dear Dockets Management Staff,
Thank you for the opportunity to comment on the discussion paper issued by the Digital Health Center of
Excellence within the Center for Devices and Radiological Health (CDRH). I submit these comments in my
personal capacity as a bilingual program evaluator and epidemiologist working in health and rehabilitation
services research in Québec, Canada, with graduate training in epidemiology, ongoing graduate studies in
business analytics, and prior clinical training in medical imaging technology. My areas of interest include
health technology assessment, health economics and outcomes research (HEOR), real-world evidence
(RWE), and causal machine learning. Institutional affiliations are provided for identification purposes only;
the views expressed are my own. I understand that this comment, including my name and contact
information, will be posted publicly on Regulations.gov.
Overall, I commend CDRH for a thoughtful, well-structured paper. The two-axis risk framework, the
competency-based premarket approach (benchmarking plus clinical confirmation), and the emphasis on
risk-proportionate postmarket monitoring across the total product life cycle (TPLC) are sound organizing
concepts that build coherently on CDRH's recent policy instruments ; the final guidance on Predetermined
Change Control Plans (PCCPs) for AI-enabled device software functions, the January 2025 draft guidance
on lifecycle management and marketing submissions for AI-enabled device software functions, and the
FDA / Health Canada / MHRA guiding principles on Good Machine Learning Practice and on transparency
for machine learning-enabled medical devices. As a Canada-based commenter, I would also underscore
the value of pursuing international alignment (including through IMDRF) as these concepts mature, so
that convergent evidence expectations reduce duplicative testing without lowering the bar for safety. My
responses to the discussion questions follow, grouped by section; where questions overlap, they are
addressed together.
Section IV: Assessment of Risk for GenAI-Enabled Devices
Question 1: Adequacy of the two-axis framework and additional dimensions.
The two axes, device activity and the consequences of relying on an incorrect output, capture the
dominant drivers of risk, and the framework is usefully continuous with the benefit-risk logic of existing
FDA guidance. I recommend keeping the framework two-dimensional and treating the additional
dimensions CDRH lists (reversibility of the resulting action, availability of downstream safeguards, time
pressure of the deployment setting, and traceability of outputs to primary sources) as documented,
auditable modifiers of a function's position on the consequences axis, rather than as additional axes; more
than two axes would compromise usability without adding discriminating power. One further modifier
warrants explicit inclusion: the detectability of an incorrect output by the intended user. An error the user
can plausibly recognize and discount (for example, a summary contradicting a visible source record) is
materially lower risk than an error the user cannot independently verify.
Question 2: The non-directive to action-directing continuum.
To give manufacturers the predictability they need, CDRH could publish a standardized directiveness
rubric with anchored levels, general information / information contextualized to the user's circumstances
/ endorsement of a specific action / specific instruction, illustrated with worked examples across clinical
contexts.
Question 3: Patient-facing versus HCP-facing functions.
Shifting patient-facing functions upward on the consequences axis is justified where the user cannot
independently evaluate the output, but the shift should be conditional on demonstrated mitigations
rather than automatic, to avoid the medical parentalism the paper rightly flags. Meaningful mitigations
include traceable sourcing to authoritative references; calibrated uncertainty language validated for lay
comprehension; conservative escalation defaults for safety-critical presentations; and reading-level and
health-literacy testing of outputs consistent with element E.4 and with human factors principles.
Demonstrated comprehension by representative lay users, including users with lower health literacy and
limited English proficiency, should be allowed to offset part of the presumptive uplift. This calibrates
protection to actual, tested user capability rather than to assumptions about it.
Question 4. Generalist versus specialist HCP-facing functions.
Three things seem proportionate: labeling that states the expected competency of the intended user
explicitly; benchmarking and human factors testing performed with adjudicators and simulated users
matching the least specialized intended user (not the most expert available); and referral-prompting
behavior, recognizing when specialist involvement is warranted, treated as a tested competency under
S.2/E.2 rather than an assumed safeguard. Where these are demonstrated, functions that extend
specialist knowledge to generalists can be a meaningful access benefit, particularly in underserved and
rural settings.
Question 5. Multi-turn conversational trajectories.
The intended use of such devices could be characterized as a behavioral envelope: the set of functions
the device is permitted to perform, the conversational conditions under which it must escalate or defer,
and the boundaries it must maintain. Premarket evidence would then demonstrate that emergent
behavior remains within the envelope across the trajectory distribution, with drift out of the envelope
(for example, informational functions migrating to action-directing behavior) treated as a reportable
failure mode rather than a labeling ambiguity.
Question 6: Under-escalation versus over-escalation.
I encourage CDRH to allow explicit decision-analytic weighting rather than a single blended accuracy
metric. Manufacturers would prespecify, per clinical context, the relative disutility of missed versus
unnecessary escalation, justified with clinical evidence and, where available, health-economic estimates
of downstream consequences; this is analogous to how screening programs weigh false negatives against
false positives. Both error rates should be reported separately with confidence intervals, and acceptable
operating points should be justified per indication, since a trade-off acceptable for pediatric sore throat
triage is not acceptable for chest pain.
Section V: Competency-Based Approach for Premarket Evaluation
Questions 7-8: Usefulness of the framework and its link to the risk framework.
The two-axis framework should drive a published, even if illustrative, evidence gradient: as a function
moves up and to the right, the breadth and difficulty of benchmarking, the independence of adjudication,
and the rigor of clinical confirmation (from retrospective evaluation toward shadow deployment and
prospective study) should increase predictably.
Question 9: Adequacy of the benchmarking elements.
Two additions merit explicit naming: (i) temporal validity, whether clinical knowledge remains current as
guidelines change, with a stated knowledge horizon and a defined update mechanism (this could be folded
into E.1 but is easily overlooked if unnamed); and (ii) input integrity and provenance handling, behavior
when inputs are incomplete, corrupted, out of distribution, or of unverifiable origin (foldable into R.1).
Under R.2, I recommend explicitly including performance in languages other than English where the
intended population includes non-English speakers; dialects and colloquialisms are named in the paper,
but cross-language performance is a distinct and empirically documented failure surface.
Question 10: Benchmark construct validity.
Sponsor-developed benchmarks are valuable for intended-use specificity and should be permitted, but
they should be complemented by at least one sequestered test set held by a party structurally
independent of both the sponsor and the foundation model developer, to guard against optimization to
the test, a role naturally suited to the third-party mechanisms discussed in Question 16.
Question 11: Selecting and justifying clinical confirmation approaches.
For real-world data and data collected outside the United States, I recommend requiring a formal
transportability assessment: characterization of case-mix, practice-pattern, coding, and data-capture
differences between source and target settings, with quantitative adjustment (for example,
standardization or reweighting) where differences are material. This is consistent with FDA's RWE
program and would be a natural application of modern causal-inference and transportability methods;
OUS evidence should be neither privileged nor discounted but transported transparently.
Questions 12-13: Statistical measurement and synthetic data.
Combining benchmarking and clinical-confirmation evidence into a single pooled performance estimate
should generally be avoided unless the sampling frames are demonstrably exchangeable; presenting them
as complementary, separately quantified lines of evidence is more defensible and more informative for
review.
Questions 14-15: Comparators and acceptance criteria.
I also support comparators reflecting the care likely to occur in the device's absence (unaided clinical
judgment, delayed specialist review, or no intervention), which is the comparator most relevant to public
health impact and to least-burdensome evaluation of access-extending devices. Usual-care
characterization from RWD and, where randomization is infeasible, target-trial-emulation designs applied
to observational data are established, principled ways to construct and justify such counterfactual
comparators.
Section VI: Postmarket Monitoring
Question 18: Accepting greater premarket uncertainty.
A pre-/post-market rebalancing is ill-suited for autonomous functions that take actions with severe
consequences; furthermore, for functions where users cannot detect errors, post-market signals would
only become apparent after harm has occurred, a situation that pre-market evidence is specifically
intended to prevent. Incorporating these conditions into the FDA’s existing guidance on benefit-risk
uncertainty would maintain a doctrinally consistent approach rather than an exceptional one.
Question 19: Postmarket performance evaluation approaches.
I encourage CDRH to add a fourth approach where feasible: outcome-linked surveillance, linking device
exposure to downstream utilization and clinical outcomes in administrative, claims, and EHR data;
behavioral metrics alone can miss harms that manifest downstream (for example, delayed care after
under-escalation). Existing infrastructure, including distributed data networks of the kind FDA already
uses for medical product surveillance, and coordinated-registry and NEST-type approaches, could be
leveraged; this is also where real-world performance monitoring expectations in the January 2025 draft
guidance connect naturally to this paper.
Questions 22-23: Modifications, re-benchmarking, and PCCPs.
Where the nature of future modifications cannot be fully prespecified, the normal condition for GenAI
systems; PCCP concepts could be adapted from prespecifying the modifications to prespecifying the
verification protocol: a “change-envelope” PCCP that fixes the competency baseline, the re-benchmarking
protocol, the acceptance criteria, and the boundaries of the authorized intended use, within which classes
of change may be implemented upon passing verification, with results documented and auditable. This
preserves the PCCP's core logic (FDA reviews the method once; the method governs many changes) while
fitting the continuous-update reality of these systems.
Section VII: Other Topics
Question 26: Agentic AI systems.
Acceptance criteria should tighten as autonomy increases and the opportunity for human review declines,
consistent with the activity axis of the risk framework, and an agentic system whose action sequences
result in control of another medical device should be evaluated against the risk profile of the controlled
device, not merely its own.
Conclusion
The discussion paper reflects a careful, risk-proportionate, least-burdensome orientation that I support.
My principal recommendations are: (i) keep the risk framework two-dimensional and operationalize the
additional dimensions, including error detectability by the user, as documented modifiers on the
consequences axis; (ii) publish an illustrative mapping from risk position to expected evidence, and an
anchored directiveness rubric applied to the empirical distribution of device outputs; (iii) require
predictive-validity evidence for gating benchmarks, with contamination testing and independent
sequestered test sets; (iv) encourage shadow deployment and require formal transportability assessment
for RWD and OUS evidence; (v) anchor subgroup performance claims in a minimum core of real data; (vi)
pair behavioral postmarket monitoring with outcome-linked surveillance in real-world data, with eventdriven re-benchmarking cadence; (vii) adapt PCCPs toward verification-protocol (“change-envelope”)
prespecification for changes that cannot be enumerated in advance; and (viii) pursue international
alignment of these concepts, including MAF content, with partner regulators and through IMDRF.
Thank you for considering these comments. I would welcome any opportunity to provide further input.
Respectfully submitted,
Sehouenou Alberic Candide Ahouehome, M.Sc., MBA
candide.ahouehome.1@ulaval.ca
Québec City, Québec, Canada