Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass
This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.
Read the source passage
Response to Discussion Questions 19, 22, and 23 The approaches described in Section VI — periodic re-benchmarking, sample-based independent clinician review, performance degradation monitoring — are all reasonable. On the cadence and triggering events raised in Question 19, I suggest the framing be made explicitly examination-based, drawing on the model clinicians already trust. Practicing clinicians do not merely have their performance monitored for drift. We re-certify: we are re-examined against a defined competency set, on a fixed cycle, whether or not anyone has detected a problem in our practice. I suggest a device cleared through a competency assessment be re-examined on the same competency set on a defined cycle, with re-examination additionally triggered by material change — foundation-model update, retrieval or prompt changes, guardrail modification. Two points follow. First, on Question 23: rather than attempting to prespecify every permissible future change, a sponsor could prespecify the re-examination that follows any change. This is a tractable commitment even where the nature of future modifications cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a modification is sensible for the clinical proficiency elements, but I would encourage CDRH to require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any change to the underlying model or guardrails, regardless of how minor the sponsor expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification examination on the grounds that little has changed in their practice, and third-party model updates are exactly the case where sponsor expectations are least reliable. 4. Foundation model MAFs: include override-relevant behavior Response to Discussion Question 25 If voluntary Foundation Model MAFs proceed, I suggest the contemplated content include, alongside architecture and training provenance: refusal behavior, content-policy changes between versions, and output stability under varied user framing — including the speaker-authority framing described in Section 1 above. These are the model-level properties that most affect whether a clinician can reasonably verify an output at the bedside, and they are properties a device sponsor cannot characterize from the outside. A sponsor cannot evaluate what the MAF does not disclose. On the incentive problem the question raises: one practical lever is that a documented Foundation Model MAF would allow sponsors to satisfy portions of the re-examination described in Section 3 above by reference, rather than by independently re-characterizing the model after every upstream update. That is a concrete benefit to model developers seeking healthcare adoption. 5. On generalizability and deployment populations Response to Discussion Questions 9 and 11 Element R.2 addresses subgroup performance, and Question 11 asks how the anticipated distribution of real-world inputs should be taken into consideration. I would encourage CDRH to treat these as one question rather than two. Recent evidence in dermatology AI indicates that distribution shift — the appearance of unfamiliar conditions — degrades performance considerably more than skin-tone differences alone [2]. Subgroup performance measured on the training-era disease mix can therefore look acceptable while real-world performance is materially worse, because what changed at deployment was the presenting case mix, not only the demographics of the patients. This bears directly on my own field. Aesthetic and dermatologic presentations in Asian populations differ substantially in disease distribution from the datasets on which most generalist models are trained, and devices cleared on North American or European evidence will encounter that shift immediately. I suggest capability assessments include test populations that differ from training populations in disease distribution as well as demographic mix, and that sponsors be asked to characterize the anticipated deployment case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset. Conclusion The physician-training analogy is the strongest idea in this paper. I encourage FDA to carry it through completely: real examinations include pressure, hierarchy, and unfamiliar patients — not only clean curricula. Devices that pass only the clean parts will fail in the clinic in exactly the ways clinicians are trained to catch, and regulators should ensure the assessment catches them first. I appreciate the opportunity to comment and am willing to provide further detail on any point. Respectfully submitted, Wen Hsien Ethan Huang, MD Founder, DrEthan AI Aesthetics ORCID 0000-0003-1727- 1870 support@drethan.ai September 4, 2026 References 1. Zhu J, et al. AI Can Be Easily Persuaded in Clinical Decision Making. arXiv:2608.29453 [preOriginal source ↗