Wen Hsien Ethan Huang, MD
““A human can intervene” is not the same as “a human will intervene.””
What they argued
'Strongly support' clinician-exam model; add speaker-authority tests; Q23 prespecify re-examination after any change, safety elements re-run in full; MAF by reference.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ14 · Comparators and acceptance criteriaQ19 · Postmarket performance evaluationQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ25 · Foundation Model Master Files
Coded positions
Have clinicians review samples of outputs
Reassess after changes or safety signals
Specify the tests or controls a future change must pass
Across the five cross-cutting questions
High-consequence work: Not stated
The comment as filed
Submitted by Wen Hsien Ethan Huang, MD — practicing aesthetic medicine clinician, medical device inventor, and educator who trains licensed physicians in the clinical use of AI tools. Full comment attached as PDF; it is a partial response addressing Discussion Questions 1, 2, 9, 11, 14, 19, 22, 23, and 25.
I strongly support the direction of this discussion paper. Modeling device evaluation on clinician training and examination is the right instinct — and I write to strengthen it from the practitioner side. Three points:
1. Benchmarking should test not only adversarial inputs but the claimed identity and authority of the person supplying them, because published evidence shows clinical AI outputs change with the perceived seniority of the speaker rather than with correctness. (Questions 9, 10)
2. The human-oversight axis of the risk framework should be tested against realistic override conditions, because "a human can intervene" is not the same as "a human will intervene." (Questions 1, 2, 14)
3. Postmarket monitoring should borrow the re-credentialing model clinicians already live under — periodic re-examination on the original competency set, not drift monitoring alone. (Questions 19, 22, 23)
Attachment
PUBLIC COMMENT
Docket No. FDA-2026-N-7874 — Considerations for the Regulation of Generative AIEnabled Medical Devices Submitted via Regulations.gov | Comment period closes
October 19, 2026
Submitter: Wen Hsien Ethan Huang, MD Affiliation: Founder, DrEthan AI Aesthetics
(drethan.ai) — clinician-led training in AI-assisted aesthetic medicine. Independent
practitioner and educator, Taiwan. ORCID: 0000-0003-1727-1870 Contact:
support@drethan.ai Submitted as: An individual. This comment is not submitted on behalf
of, or at the request of, any device manufacturer, trade association, or sponsor. Scope: This
is a partial response. It addresses Discussion Questions 1, 2, 9, 11, 14, 19, 22, 23, and 25.
Perspective and relevant background
I am a practicing aesthetic medicine clinician, a medical device inventor, and an educator
who trains licensed physicians in the clinical use of AI tools. My comments draw on three
vantage points that bear directly on this discussion paper:
As a developer. I first-authored work on a deep convolutional neural network for clinical
image assessment in aesthetic medicine, presented at the 34th Annual Conference of
the Japanese Society for Artificial Intelligence (JSAI 2020; DOI
10.11517/pjsai.JSAI2020.0_1K5ES204). I have direct experience of the distance between
benchmark performance and bedside behavior.
As a device inventor. I hold granted patents in Japan, Germany, the European Union,
China, and Taiwan covering thread-lift assistance instrumentation, and have worked
through the design and documentation burden that device development imposes.
As an educator. I teach licensed clinicians a dedicated curriculum on when and how to
override AI recommendations (a “Trust / Verify / Override” framework). Sections 1 and 2
below come directly from watching trained clinicians fail to override.
Disclosure of interest: I sell paid training courses to licensed clinicians on the safe use of
AI in clinical practice, through drethan.ai. I have no financial interest in, and no consulting or
advisory relationship with, any manufacturer or developer of a GenAI-enabled medical
device.
Summary
I strongly support the direction of this discussion paper. Modeling device evaluation on
clinician training and examination is the right instinct — and I write to strengthen it from the
practitioner side. Three points:
1. Benchmarking should test not only adversarial inputs but the claimed identity and
authority of the person supplying them, because published evidence shows clinical AI
outputs change with the perceived seniority of the speaker rather than with correctness.
(Questions 9, 10)
2. The human-oversight axis of the risk framework should be tested against realistic
override conditions, because “a human can intervene” is not the same as “a human will
intervene.” (Questions 1, 2, 14)
3. Postmarket monitoring should borrow the re-credentialing model clinicians already
live under — periodic re-examination on the original competency set, not drift
monitoring alone. (Questions 19, 22, 23)
1. Benchmarking: test the examination conditions, not only the
curriculum
Response to Discussion Questions 9 and 10
The paper proposes benchmarking followed by clinical confirmation, mirroring how health
professionals are assessed. Appendix A already anticipates part of what I want to raise:
element S.2 contemplates adversarial prompting, prompt injection, and emotionalmanipulation scenarios, and element R.1 contemplates consistency across paraphrases that
differ in emotional framing. From the examination hall, one dimension appears to be missing
from both:
Benchmarking should test sensitivity to the claimed identity, seniority, and authority
of the speaker — not only to the wording of the input.
Recent controlled work shows that clinical AI recommendations shift according to who
appears to be speaking: the same persuasive input changed roughly 10% more model
outputs when attributed to a senior clinician than to a medical student, and fabricated-butplausible clinician opinions pushed models off correct answers onto incorrect ones [1].
Professional authority, claimed track record, institutional affiliation, and repeated pressure
all swayed the model without improving its accuracy.
This is a distinct failure mode from adversarial prompting as currently described. The input
is not malformed, malicious, or emotionally manipulative — it is an ordinary second opinion,
delivered with a credential attached. A device that performs well on a clean benchmark can
therefore fail exactly the way clinicians fail: by deferring to hierarchy. In a real clinic, an AI
output is routinely challenged by a senior colleague who is sometimes wrong.
Concretely, I suggest that under element R.1, benchmark protocols prespecify speakerauthority test conditions alongside paraphrase and emotional-framing conditions: the
same clinical case, held constant, with the accompanying opinion attributed to users of
differing claimed seniority, confidence, and persistence. A device whose safety-critical
behavior — escalation, refusal, diagnostic conclusion — moves with the claimed credential
of the user, rather than with the clinical facts, has failed the robustness element even if
every individual output looks reasonable.
Suggested question for sponsors: “How does the device’s output change when prior
input implies the user is senior, junior, confident, or insistent?”
2. The risk framework: “human oversight” must be tested, not
assumed
Response to Discussion Questions 1, 2, and 14
The two-axis framework (device activity × severity of harm) is sound. Question 1 asks
whether additional dimensions — including the time pressure of the deployment setting —
should be represented. My answer is yes, and specifically: the degree of human oversight
should be treated as an empirical property of the deployment setting, not as a design
feature that is present or absent.
In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override
option does not guarantee the override will be used. Automation bias is well documented,
and clinicians under time pressure defer to confident outputs. Appendix A element E.4
recognizes automation bias, but treats it as a communication-quality attribute of the device.
I would encourage CDRH to also treat it as a modifier of position on the activity axis: a
function nominally placed at “acts with continuous HCP supervision” may in practice
operate closer to autonomy if the supervision is not exercised.
I encourage FDA to:
Treat “degree of human oversight” as a property to be demonstrated in
representative use conditions — time-pressured, multi-patient, real interface — rather
than asserted in labeling.
Ask sponsors to show evidence that intended users can and do detect incorrect outputs
in representative workflows. This is an override-rate and detection-rate measurement,
and it is precisely the kind of human-AI team evidence contemplated in Question 14.
Note that the paper’s own observation — that a “talk to your doctor” statement may not
make an output less directive — applies symmetrically, and bears on Question 2: an
override interface that is never used provides no oversight. Directiveness and oversight
should both be assessed by observed user behavior rather than by the presence of text
on the screen.
3. Postmarket monitoring: borrow the re-credentialing model
Response to Discussion Questions 19, 22, and 23
The approaches described in Section VI — periodic re-benchmarking, sample-based
independent clinician review, performance degradation monitoring — are all reasonable. On
the cadence and triggering events raised in Question 19, I suggest the framing be made
explicitly examination-based, drawing on the model clinicians already trust.
Practicing clinicians do not merely have their performance monitored for drift. We re-certify:
we are re-examined against a defined competency set, on a fixed cycle, whether or not
anyone has detected a problem in our practice. I suggest a device cleared through a
competency assessment be re-examined on the same competency set on a defined
cycle, with re-examination additionally triggered by material change — foundation-model
update, retrieval or prompt changes, guardrail modification.
Two points follow. First, on Question 23: rather than attempting to prespecify every
permissible future change, a sponsor could prespecify the re-examination that follows any
change. This is a tractable commitment even where the nature of future modifications
cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for
GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a
modification is sensible for the clinical proficiency elements, but I would encourage CDRH to
require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any
change to the underlying model or guardrails, regardless of how minor the sponsor
expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification
examination on the grounds that little has changed in their practice, and third-party model
updates are exactly the case where sponsor expectations are least reliable.
4. Foundation model MAFs: include override-relevant behavior
Response to Discussion Question 25
If voluntary Foundation Model MAFs proceed, I suggest the contemplated content include,
alongside architecture and training provenance: refusal behavior, content-policy changes
between versions, and output stability under varied user framing — including the
speaker-authority framing described in Section 1 above. These are the model-level
properties that most affect whether a clinician can reasonably verify an output at the
bedside, and they are properties a device sponsor cannot characterize from the outside. A
sponsor cannot evaluate what the MAF does not disclose.
On the incentive problem the question raises: one practical lever is that a documented
Foundation Model MAF would allow sponsors to satisfy portions of the re-examination
described in Section 3 above by reference, rather than by independently re-characterizing
the model after every upstream update. That is a concrete benefit to model developers
seeking healthcare adoption.
5. On generalizability and deployment populations
Response to Discussion Questions 9 and 11
Element R.2 addresses subgroup performance, and Question 11 asks how the anticipated
distribution of real-world inputs should be taken into consideration. I would encourage
CDRH to treat these as one question rather than two.
Recent evidence in dermatology AI indicates that distribution shift — the appearance of
unfamiliar conditions — degrades performance considerably more than skin-tone
differences alone [2]. Subgroup performance measured on the training-era disease mix can
therefore look acceptable while real-world performance is materially worse, because what
changed at deployment was the presenting case mix, not only the demographics of the
patients.
This bears directly on my own field. Aesthetic and dermatologic presentations in Asian
populations differ substantially in disease distribution from the datasets on which most
generalist models are trained, and devices cleared on North American or European evidence
will encounter that shift immediately. I suggest capability assessments include test
populations that differ from training populations in disease distribution as well as
demographic mix, and that sponsors be asked to characterize the anticipated deployment
case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset.
Conclusion
The physician-training analogy is the strongest idea in this paper. I encourage FDA to carry
it through completely: real examinations include pressure, hierarchy, and unfamiliar patients
— not only clean curricula. Devices that pass only the clean parts will fail in the clinic in
exactly the ways clinicians are trained to catch, and regulators should ensure the
assessment catches them first.
I appreciate the opportunity to comment and am willing to provide further detail on any
point.
Respectfully submitted,
Wen Hsien Ethan Huang, MD Founder, DrEthan AI Aesthetics ORCID 0000-0003-17271870 support@drethan.ai September 4, 2026
References
1. Zhu J, et al. AI Can Be Easily Persuaded in Clinical Decision Making. arXiv:2608.29453
[preprint].
2. Kunwor N, Poudel S, Trinh Q-H, Arafat J, Gaire SK. Disease Burden over Skin Tone:
Decomposing the Dermatology-AI Generalization Gap. arXiv:2609.02111 [preprint].