← All 95 filings

Wen Hsien Ethan Huang, MD

CliniciansClinicianFiled September 3, 20262,030 words · 1 attachmentFDA-2026-N-7874-0057
““A human can intervene” is not the same as “a human will intervene.””

What they argued

RecovryAI’s one-line reading of the filing.

'Strongly support' clinician-exam model; add speaker-authority tests; Q23 prespecify re-examination after any change, safety elements re-run in full; MAF by reference.

Themes it raises

8 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“My answer is yes, and specifically: the degree of human oversight should be treated as an empirical property of the deployment setting, not as a design feature that is present or absent.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“The physician-training analogy is the strongest idea in this paper.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Benchmarking should test sensitivity to the claimed identity, seniority, and authority of the speaker — not only to the wording of the input.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“I suggest capability assessments include test populations that differ from training populations in disease distribution as well as demographic mix, and that sponsors be asked to characterize the anticipated deployment case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset.”
Watching the device after it shipsFDA Q19, Q20
“I suggest a device cleared through a competency assessment be re-examined on the same competency set on a defined cycle, with re-examination additionally triggered by material change — foundation-model update, retrieval or prompt changes, guardrail modification.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“First, on Question 23: rather than attempting to prespecify every permissible future change, a sponsor could prespecify the re-examination that follows any change.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override option does not guarantee the override will be used.”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“Aesthetic and dermatologic presentations in Asian populations differ substantially in disease distribution from the datasets on which most generalist models are trained, and devices cleared on North American or European evidence will encounter that shift immediately.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ14 · Comparators and acceptance criteriaQ19 · Postmarket performance evaluationQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ25 · Foundation Model Master Files

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
Keep it, but add or change elements
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider the user and clinical context
Q11When can a device be confirmed without a prospective clinical study, and what earns that lighter path?
Check that evidence fits the intended users and setting
Q14For open-ended AI outputs, who is the performance comparator: a clinician panel, generalists, specialists, or the human-AI team?
Evaluate the clinician and AI working together
Q19How should an AI device be monitored after launch, and what sets the cadence?
Repeat performance testing on a schedule
Have clinicians review samples of outputs
Reassess after changes or safety signals
Q22With the premarket competency assessment as the baseline, which post-deployment changes need re-evaluation, and how much?
Scale retesting to the change’s clinical impact
Q23How can a change-control plan cover changes that cannot be fully specified in advance?
Define what must remain safe instead of predicting every edit
Specify the tests or controls a future change must pass
Q25Would voluntary Foundation Model Master Files be practical, and useful in premarket review?
Use them if specified conditions are met

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
No position stated
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
No position stated
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Not stated
High-consequence work: Not stated
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Submitted by Wen Hsien Ethan Huang, MD — practicing aesthetic medicine clinician, medical device inventor, and educator who trains licensed physicians in the clinical use of AI tools. Full comment attached as PDF; it is a partial response addressing Discussion Questions 1, 2, 9, 11, 14, 19, 22, 23, and 25.

I strongly support the direction of this discussion paper. Modeling device evaluation on clinician training and examination is the right instinct — and I write to strengthen it from the practitioner side. Three points:

1. Benchmarking should test not only adversarial inputs but the claimed identity and authority of the person supplying them, because published evidence shows clinical AI outputs change with the perceived seniority of the speaker rather than with correctness. (Questions 9, 10)

2. The human-oversight axis of the risk framework should be tested against realistic override conditions, because "a human can intervene" is not the same as "a human will intervene." (Questions 1, 2, 14)

3. Postmarket monitoring should borrow the re-credentialing model clinicians already live under — periodic re-examination on the original competency set, not drift monitoring alone. (Questions 19, 22, 23)

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

PUBLIC COMMENT
Docket No. FDA-2026-N-7874 — Considerations for the Regulation of Generative AIEnabled Medical Devices Submitted via Regulations.gov | Comment period closes
October 19, 2026

Submitter: Wen Hsien Ethan Huang, MD Affiliation: Founder, DrEthan AI Aesthetics
(drethan.ai) — clinician-led training in AI-assisted aesthetic medicine. Independent
practitioner and educator, Taiwan. ORCID: 0000-0003-1727-1870 Contact:
support@drethan.ai Submitted as: An individual. This comment is not submitted on behalf
of, or at the request of, any device manufacturer, trade association, or sponsor. Scope: This
is a partial response. It addresses Discussion Questions 1, 2, 9, 11, 14, 19, 22, 23, and 25.

Perspective and relevant background
I am a practicing aesthetic medicine clinician, a medical device inventor, and an educator
who trains licensed physicians in the clinical use of AI tools. My comments draw on three
vantage points that bear directly on this discussion paper:

As a developer. I first-authored work on a deep convolutional neural network for clinical
image assessment in aesthetic medicine, presented at the 34th Annual Conference of
the Japanese Society for Artificial Intelligence (JSAI 2020; DOI
10.11517/pjsai.JSAI2020.0_1K5ES204). I have direct experience of the distance between
benchmark performance and bedside behavior.

As a device inventor. I hold granted patents in Japan, Germany, the European Union,
China, and Taiwan covering thread-lift assistance instrumentation, and have worked
through the design and documentation burden that device development imposes.

As an educator. I teach licensed clinicians a dedicated curriculum on when and how to
override AI recommendations (a “Trust / Verify / Override” framework). Sections 1 and 2
below come directly from watching trained clinicians fail to override.

Disclosure of interest: I sell paid training courses to licensed clinicians on the safe use of
AI in clinical practice, through drethan.ai. I have no financial interest in, and no consulting or
advisory relationship with, any manufacturer or developer of a GenAI-enabled medical
device.
Summary
I strongly support the direction of this discussion paper. Modeling device evaluation on
clinician training and examination is the right instinct — and I write to strengthen it from the
practitioner side. Three points:

1. Benchmarking should test not only adversarial inputs but the claimed identity and
authority of the person supplying them, because published evidence shows clinical AI
outputs change with the perceived seniority of the speaker rather than with correctness.
(Questions 9, 10)

2. The human-oversight axis of the risk framework should be tested against realistic
override conditions, because “a human can intervene” is not the same as “a human will
intervene.” (Questions 1, 2, 14)

3. Postmarket monitoring should borrow the re-credentialing model clinicians already
live under — periodic re-examination on the original competency set, not drift
monitoring alone. (Questions 19, 22, 23)

1. Benchmarking: test the examination conditions, not only the
curriculum
Response to Discussion Questions 9 and 10

The paper proposes benchmarking followed by clinical confirmation, mirroring how health
professionals are assessed. Appendix A already anticipates part of what I want to raise:
element S.2 contemplates adversarial prompting, prompt injection, and emotionalmanipulation scenarios, and element R.1 contemplates consistency across paraphrases that
differ in emotional framing. From the examination hall, one dimension appears to be missing
from both:

Benchmarking should test sensitivity to the claimed identity, seniority, and authority
of the speaker — not only to the wording of the input.

Recent controlled work shows that clinical AI recommendations shift according to who
appears to be speaking: the same persuasive input changed roughly 10% more model
outputs when attributed to a senior clinician than to a medical student, and fabricated-butplausible clinician opinions pushed models off correct answers onto incorrect ones [1].
Professional authority, claimed track record, institutional affiliation, and repeated pressure
all swayed the model without improving its accuracy.

This is a distinct failure mode from adversarial prompting as currently described. The input
is not malformed, malicious, or emotionally manipulative — it is an ordinary second opinion,
delivered with a credential attached. A device that performs well on a clean benchmark can
therefore fail exactly the way clinicians fail: by deferring to hierarchy. In a real clinic, an AI
output is routinely challenged by a senior colleague who is sometimes wrong.

Concretely, I suggest that under element R.1, benchmark protocols prespecify speakerauthority test conditions alongside paraphrase and emotional-framing conditions: the
same clinical case, held constant, with the accompanying opinion attributed to users of
differing claimed seniority, confidence, and persistence. A device whose safety-critical
behavior — escalation, refusal, diagnostic conclusion — moves with the claimed credential
of the user, rather than with the clinical facts, has failed the robustness element even if
every individual output looks reasonable.

Suggested question for sponsors: “How does the device’s output change when prior
input implies the user is senior, junior, confident, or insistent?”

2. The risk framework: “human oversight” must be tested, not
assumed
Response to Discussion Questions 1, 2, and 14

The two-axis framework (device activity × severity of harm) is sound. Question 1 asks
whether additional dimensions — including the time pressure of the deployment setting —
should be represented. My answer is yes, and specifically: the degree of human oversight
should be treated as an empirical property of the deployment setting, not as a design
feature that is present or absent.

In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override
option does not guarantee the override will be used.
Automation bias is well documented,
and clinicians under time pressure defer to confident outputs. Appendix A element E.4
recognizes automation bias, but treats it as a communication-quality attribute of the device.
I would encourage CDRH to also treat it as a modifier of position on the activity axis: a
function nominally placed at “acts with continuous HCP supervision” may in practice
operate closer to autonomy if the supervision is not exercised.

I encourage FDA to:

Treat “degree of human oversight” as a property to be demonstrated in
representative use conditions — time-pressured, multi-patient, real interface — rather
than asserted in labeling.

Ask sponsors to show evidence that intended users can and do detect incorrect outputs
in representative workflows. This is an override-rate and detection-rate measurement,
and it is precisely the kind of human-AI team evidence contemplated in Question 14.

Note that the paper’s own observation — that a “talk to your doctor” statement may not
make an output less directive — applies symmetrically, and bears on Question 2: an
override interface that is never used provides no oversight. Directiveness and oversight
should both be assessed by observed user behavior rather than by the presence of text
on the screen.

3. Postmarket monitoring: borrow the re-credentialing model
Response to Discussion Questions 19, 22, and 23

The approaches described in Section VI — periodic re-benchmarking, sample-based
independent clinician review, performance degradation monitoring — are all reasonable. On
the cadence and triggering events raised in Question 19, I suggest the framing be made
explicitly examination-based, drawing on the model clinicians already trust.

Practicing clinicians do not merely have their performance monitored for drift. We re-certify:
we are re-examined against a defined competency set, on a fixed cycle, whether or not
anyone has detected a problem in our practice. I suggest a device cleared through a
competency assessment be re-examined on the same competency set on a defined
cycle, with re-examination additionally triggered by material change — foundation-model
update, retrieval or prompt changes, guardrail modification.

Two points follow. First, on Question 23: rather than attempting to prespecify every
permissible future change, a sponsor could prespecify the re-examination that follows any
change.
This is a tractable commitment even where the nature of future modifications
cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for
GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a
modification is sensible for the clinical proficiency elements, but I would encourage CDRH to
require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any
change to the underlying model or guardrails, regardless of how minor the sponsor
expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification
examination on the grounds that little has changed in their practice, and third-party model
updates are exactly the case where sponsor expectations are least reliable.

4. Foundation model MAFs: include override-relevant behavior
Response to Discussion Question 25

If voluntary Foundation Model MAFs proceed, I suggest the contemplated content include,
alongside architecture and training provenance: refusal behavior, content-policy changes
between versions, and output stability under varied user framing — including the
speaker-authority framing described in Section 1 above. These are the model-level
properties that most affect whether a clinician can reasonably verify an output at the
bedside, and they are properties a device sponsor cannot characterize from the outside. A
sponsor cannot evaluate what the MAF does not disclose.

On the incentive problem the question raises: one practical lever is that a documented
Foundation Model MAF would allow sponsors to satisfy portions of the re-examination
described in Section 3 above by reference, rather than by independently re-characterizing
the model after every upstream update. That is a concrete benefit to model developers
seeking healthcare adoption.

5. On generalizability and deployment populations
Response to Discussion Questions 9 and 11

Element R.2 addresses subgroup performance, and Question 11 asks how the anticipated
distribution of real-world inputs should be taken into consideration. I would encourage
CDRH to treat these as one question rather than two.

Recent evidence in dermatology AI indicates that distribution shift — the appearance of
unfamiliar conditions — degrades performance considerably more than skin-tone
differences alone [2]. Subgroup performance measured on the training-era disease mix can
therefore look acceptable while real-world performance is materially worse, because what
changed at deployment was the presenting case mix, not only the demographics of the
patients.

This bears directly on my own field. Aesthetic and dermatologic presentations in Asian
populations differ substantially in disease distribution from the datasets on which most
generalist models are trained, and devices cleared on North American or European evidence
will encounter that shift immediately.
I suggest capability assessments include test
populations that differ from training populations in disease distribution as well as
demographic mix, and that sponsors be asked to characterize the anticipated deployment
case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset.

Conclusion
The physician-training analogy is the strongest idea in this paper. I encourage FDA to carry
it through completely: real examinations include pressure, hierarchy, and unfamiliar patients
— not only clean curricula. Devices that pass only the clean parts will fail in the clinic in
exactly the ways clinicians are trained to catch, and regulators should ensure the
assessment catches them first.

I appreciate the opportunity to comment and am willing to provide further detail on any
point.

Respectfully submitted,

Wen Hsien Ethan Huang, MD Founder, DrEthan AI Aesthetics ORCID 0000-0003-17271870 support@drethan.ai September 4, 2026

References
1. Zhu J, et al. AI Can Be Easily Persuaded in Clinical Decision Making. arXiv:2608.29453
[preprint].

2. Kunwor N, Poudel S, Trinh Q-H, Arafat J, Gaire SK. Disease Burden over Skin Tone:
Decomposing the Dermatology-AI Generalization Gap. arXiv:2609.02111 [preprint].