FDA GenAI discussion / Question 23 of 26

How can a change-control plan cover changes that cannot be fully specified in advance?

Full FDA question

How might PCCP concepts or other change-control approaches be adapted for GenAI-enabled devices when the nature or scope of future modifications cannot be fully prespecified?
Read the FDA discussion paper ↗

21 of 95 submissions reference this question.

All audiences
13 Industry4 Clinicians1 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/23
Filter by audience
Question 23 · Public feedback

What respondents recommend

9 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
Response to Discussion Questions 19, 22, and 23 The approaches described in Section VI — periodic re-benchmarking, sample-based independent clinician review, performance degradation monitoring — are all reasonable. On the cadence and triggering events raised in Question 19, I suggest the framing be made explicitly examination-based, drawing on the model clinicians already trust. Practicing clinicians do not merely have their performance monitored for drift. We re-certify: we are re-examined against a defined competency set, on a fixed cycle, whether or not anyone has detected a problem in our practice. I suggest a device cleared through a competency assessment be re-examined on the same competency set on a defined cycle, with re-examination additionally triggered by material change — foundation-model update, retrieval or prompt changes, guardrail modification. Two points follow. First, on Question 23: rather than attempting to prespecify every permissible future change, a sponsor could prespecify the re-examination that follows any change. This is a tractable commitment even where the nature of future modifications cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a modification is sensible for the clinical proficiency elements, but I would encourage CDRH to require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any change to the underlying model or guardrails, regardless of how minor the sponsor expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification examination on the grounds that little has changed in their practice, and third-party model updates are exactly the case where sponsor expectations are least reliable. 4. Foundation model MAFs: include override-relevant behavior Response to Discussion Question 25 If voluntary Foundation Model MAFs proceed, I suggest the contemplated content include, alongside architecture and training provenance: refusal behavior, content-policy changes between versions, and output stability under varied user framing — including the speaker-authority framing described in Section 1 above. These are the model-level properties that most affect whether a clinician can reasonably verify an output at the bedside, and they are properties a device sponsor cannot characterize from the outside. A sponsor cannot evaluate what the MAF does not disclose. On the incentive problem the question raises: one practical lever is that a documented Foundation Model MAF would allow sponsors to satisfy portions of the re-examination described in Section 3 above by reference, rather than by independently re-characterizing the model after every upstream update. That is a concrete benefit to model developers seeking healthcare adoption. 5. On generalizability and deployment populations Response to Discussion Questions 9 and 11 Element R.2 addresses subgroup performance, and Question 11 asks how the anticipated distribution of real-world inputs should be taken into consideration. I would encourage CDRH to treat these as one question rather than two. Recent evidence in dermatology AI indicates that distribution shift — the appearance of unfamiliar conditions — degrades performance considerably more than skin-tone differences alone [2]. Subgroup performance measured on the training-era disease mix can therefore look acceptable while real-world performance is materially worse, because what changed at deployment was the presenting case mix, not only the demographics of the patients. This bears directly on my own field. Aesthetic and dermatologic presentations in Asian populations differ substantially in disease distribution from the datasets on which most generalist models are trained, and devices cleared on North American or European evidence will encounter that shift immediately. I suggest capability assessments include test populations that differ from training populations in disease distribution as well as demographic mix, and that sponsors be asked to characterize the anticipated deployment case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset. Conclusion The physician-training analogy is the strongest idea in this paper. I encourage FDA to carry it through completely: real examinations include pressure, hierarchy, and unfamiliar patients — not only clean curricula. Devices that pass only the clean parts will fail in the clinic in exactly the ways clinicians are trained to catch, and regulators should ensure the assessment catches them first. I appreciate the opportunity to comment and am willing to provide further detail on any point. Respectfully submitted, Wen Hsien Ethan Huang, MD Founder, DrEthan AI Aesthetics ORCID 0000-0003-1727- 1870 support@drethan.ai September 4, 2026 References 1. Zhu J, et al. AI Can Be Easily Persuaded in Clinical Decision Making. arXiv:2608.29453 [pre
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
Discussion Question 23 – PCCP Concepts and their Use for Postmarket Monitoring PCCPs are an important approach for AI-enabled devices and it is reasonable to extend this concept to GenAI- enabled systems. However, it should be recognized that the exact nature of all future modifications may not be fully predictable. For GenAI, PCCPs may therefore need to rely more heavily on procedural and boundary-based prespecification rather than exhaustive technical prespecification of every future modification. For example, this applies to the assumptions about application environment, user interaction, or safeguards, as defined during Criticality Assessment or addressed during Product-Specific Risk Management. The competency-based approach could provide a useful but extended structure for such PCCPs. In comparison to currently pursued PCCP approaches, the stages from restricted or closely supervised use toward greater operational independence may be integrated into this approach. Progression between these stages should occur only when predefined competency and real-world performance criteria have been met. Conversely, emerging postmarket signals may lead to a downgrading, such as reducing scope or autonomy or increasing required human supervision. PCCPs should also account for changes in the environment and the continued validity of the evaluation framework. A device may remain technically unchanged while new clinical guidelines, workflows, user competencies, or other environmental factors invalidate assumptions underlying its prior assessment.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass · Define when a change needs further review

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass; define when a change needs further review. The passage gives the applicable scope and conditions.

Read the source passage
Question 23 - PCCPs when future modifications cannot be fully prespecified When exact future modifications cannot be known, PCCP concepts can still be useful if the sponsor can prespecify classes or bounds of change, detection mechanisms, evaluation methodology, acceptance criteria, monitoring requirements, rollback conditions, and escalation thresholds. The emphasis may appropriately shift from prespecifying every future parameter value to prespecifying the governance process by which a bounded class of modifications will be detected, qualified, tested, and either qualified for use, restricted, subjected to additional review, rolled back, or prevented from entering the affected clinical function. For example, a sponsor might prespecify a class of retrieval-source updates limited to approved clinical knowledge repositories, together with defined validation tests, acceptance thresholds, rollback criteria, and escalation requirements, even if the exact future source update cannot be identified in advance. Whether and to what extent such an approach can be accommodated within current PCCP authorities, or would require additional policy development, is a separate legal and regulatory question. [2] 12
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
Question 23 — PCCPs and unforeseeable future modifications A PCCP should define the fence, not predict every future move inside it. It is unrealistic to pre-specify every future change to prompts, retrieval, orchestration, or model weights. A PCCP should instead state what the system is authorized to do, what performance it must maintain, what boundaries cannot be crossed, and what triggers automatic revalidation — regulating the boundaries of acceptable evolution rather than attempting to forecast it.
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Specify the tests or controls a future change must pass

This filing recommends: specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
1. Compression of the model artifact should be named as a change category (Questions 22, 23, 19) Section VI.C enumerates three kinds of post-deployment change: sponsor-initiated discrete modifications such as software updates, algorithm revisions, retraining events, and changes to intended functionality; model-evolution changes occurring passively as the device adapts during use; and unplanned changes arising from updates to a third-party foundation model. Compression of the model artifact fits none of these cleanly. Quantization, pruning, and distillation reduce the memory footprint and hardware requirements of a model (see, e.g., Frantar et al., arXiv:2210.17323; Dettmers et al., arXiv:2305.14314). For devices intended to run on constrained hardware, at the point of care, or inside controlled network environments, compression is frequently not optional. Small models deployed on premises or at the edge often require quantization to meet deployment constraints, and for many such deployments compression is the step that makes the deployment feasible at all. Compression is typically applied late, after functional validation, and often by an infrastructure or deployment function rather than the team responsible for safety assessment. It is not a retraining event, it is not passive model evolution, and it is not initiated by an upstream developer. Under the taxonomy as written, a sponsor acting in good faith could conclude that recompressing a model to fit a new hardware target falls within no enumerated change category. Question 22 asks whether there are categories of change that might not significantly affect safety or effectiveness and might be managed within a sponsor's quality management system rather than requiring premarket review. I expect compression to be nominated for that treatment, and would urge caution. Compression alters the numerical representation of model weights, and therefore alters output distributions at the token-probability boundaries where refusal and compliance decisions are resolved. Whether that alteration is behaviorally material is an empirical question whose answer varies by method, precision, model family, and behavior measured. Reported findings in the literature diverge, and the divergence is itself the point: compression is not a single operation, the methods do not behave equivalently, and the direction of effect is not predictable in advance from the compression parameters alone. The structural difficulty is that a sponsor evaluating only one artifact cannot distinguish between these possibilities. Evaluate only the full-precision checkpoint and the results may not describe what the device does. Evaluate only the compressed artifact and the results are representative of the device but cannot attribute behavior between the base model and the compression step, which matters when the base model is later updated Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 2 or the precision target changes. Naming compression as a change category is what forces the comparison that resolves the ambiguity in either direction. Recommendations: • Add compression of the model artifact, including quantization, pruning, and distillation, to the enumerated change categories in Section VI.C. • A sponsor should validate the exact deployed artifact after a material compression or inference-stack change, where the artifact is understood to include numerical precision and compression method, inference runtime and accelerator class, decoding configuration, system prompt, retrieval pipeline, and tool policy. Evidence may be risk-proportionate, but should include targeted regression testing of the Safety and Generalizability elements rather than general performance evaluation alone. • If compression is considered for inclusion in a PCCP, the plan should prespecify the compression method, precision target, and re-benchmarking evidence, rather than treating compression as a presumptively low-impact class of change. 2. Representative real-world sampling is not designed to estimate adversarial resistance (Question 19) Section VI.A describes three postmarket approaches: periodic device benchmarking, periodic sample-based clinician review, and performance degradation monitoring. My comment concerns what the second and third can and cannot support. Sample-based clinician review draws on real-world inputs and outputs, sampled to reflect the range of clinically relevant presentations the device encounters. This is well suited to detecting degradation in clinical proficiency, the E-series elements. Unless it is deliberately supplemented with adversarially constructed probes, representative real-world sampling is not designed to estimate, and should not be treated as evidence of, resistance to prompt injection, adversarial scope testing, emotional manipulation, or multi-turn escalation. Sampling designed to represent the distribution of clinical presentations will not contain these inputs at a
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
Question 23 — How should FDA address future AI modifications that cannot be fully predicted in advance? Response If future changes cannot be predicted precisely, regulators should define the safety properties that must remain invariant even as the technology evolves. Those safety invariants should include, where relevant:  grounding in current evidence;  appropriate clinical escalation;  preservation of human override;  auditability;  provenance;  appropriate uncertainty communication; Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 22 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  reliable detection of safety-critical conditions;  maintenance of validated level-of-care logic;  and continued ability to reconstruct consequential decisions after the fact. A manufacturer may not be able to predict every future capability. It should nevertheless be able to say: “Regardless of how the system evolves, these safety constraints cannot disappear without revalidation.” That is especially important for adaptive or agentic systems whose future behaviors may emerge from combinations of tools and capabilities that were not fully anticipated at initial approval.
Original source ↗

Nathan Sabich

Industry · Aug 18, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass. The passage gives the applicable scope and conditions.

Read the source passage
FDA Question 23: How might PCCP concepts or other change-control approaches be adapted for GenAI-enabled devices when the nature or scope of future modifications cannot be fully prespecified? Opinion: The PCCP Framework requires pre-specification of modifications, but GenAI foundation models evolve in ways that cannot be fully predicted. The solution is not to abandon pre-specification but to pre-specify at a higher-level of abstraction. Rather than specifying “the model will be retrained with X additional images,” the PCCP for a GenAI-enabled device should specify that the foundation model may be updated provided the updated model meets the following competency benchmarks, passes the following safety evaluations, and does not exceed the following performance deviation bounds from the authorized baseline. Architecture to Address This: I have already created for an architecture's governance framework that operates at exactly this level of abstraction. The AI lifecycle management controls apply across all the steps I created in my architecture in order to provide continuous QMS oversight (design controls, CAPA, change management, documentation), risk management (ISO 14971 aligned), human oversight (HITL throughout with escalation paths), cybersecurity (secure by design, threat monitoring, incident response), and transparency and labeling (intended use, limitations, performance, updates). This governance layer is change-agnostic and it applies regardless of whether the modification is a retraining event, a foundation model version update, or a prompt engineering adjustment. The key is that the governance controls remain constant even as the specific modifications vary. Recommendation to FDA: Create a new PCCP category called "Performance-Bounded PCCP" for GenAI-enabled devices. Unlike a traditional PCCP that pre-specifies the exact modifications, a "Performance-Bounded PCCP" would pre-specify: (1) the competency benchmarks the device must continue to meet after any modification, (2) the maximum acceptable performance deviation from the authorized baseline across each benchmark dimension, (3) the monitoring methodology for detecting deviations, (4) the rollback protocol if deviations are detected, and (5) the cumulative impact tracking methodology. Any modification that keeps the device within the performance bounds may be implemented without a new submission. Any modification that causes the device to fall outside the bounds triggers either a rollback or a new submission. This approach accommodates the unpredictability of GenAI evolution while maintaining the principle that FDA authorizes the safety and effectiveness envelope, not the specific technical implementation.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass · Define when a change needs further review

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass; define when a change needs further review. The passage gives the applicable scope and conditions.

Read the source passage
Question 23 — PCCPs when modifications cannot be prespecified This is the right question and I think it admits a clean answer, which is the most useful thing I can offer in this comment. 14 of 19 Docket No. FDA-2026-N-7874 The current PCCP construct prespecifies what will change — the modification, the methods to implement and validate it, the impact assessment. For generative devices this fails, not because sponsors are unwilling but because the change space is genuinely open: a sponsor cannot enumerate in advance the model versions a provider will release or the prompt refinements that experience will suggest. The adaptation is to invert what is locked. An invariant-based PCCP prespecifies not the change but what must remain true after any change: • The evaluation instrument — the complete benchmark suite, escrowed and immutable. • The acceptance thresholds, including subgroup-level thresholds and the non-inferiority margin. • The verification protocol — what is run, in what order, by whom, with what independence, before any change reaches users. • The rollback criteria and mechanism, including maximum time-to-rollback and the technical demonstration that rollback is achievable. • The change categories in scope, defined by the Tier 2 boundary above rather than by enumeration of specific changes. • The documentation and notification obligations attaching to each change. Under this structure a sponsor may make any change falling within the scoped categories, provided the changed device passes the locked evaluation at the locked thresholds under the locked protocol, with everything documented and inspectable. The sponsor gains the operational flexibility the technology requires. CDRH retains a fixed, auditable gate that does not move. The construct depends entirely on the integrity of the locked instrument. If the sponsor can modify the benchmark suite while modifying the device, the whole thing is circular and provides no assurance at all. Escrow of the suite, and Tier 3 treatment of any change to it, are not optional features of this proposal — they are the proposal.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Define what must remain safe instead of predicting every edit · Specify the tests or controls a future change must pass · Define when a change needs further review

This filing recommends: define what must remain safe instead of predicting every edit; specify the tests or controls a future change must pass; define when a change needs further review. The passage gives the applicable scope and conditions.

Read the source passage
FDA Question 23 - PCCP when future changes not fully known Trace ID. TR-Q23 | FDA Q23; Sec. VI.D; App. B; pp. 21-22 / 29-30 BCR response. Prespecify an allowable change envelope rather than every exact future edit: which boundaries may move, maximum deltas, required tests, hold points, rollback conditions, and new-submission triggers. BCR rule basis. BCR-R02,R09,R15,R16,R17 Solution-stack link. S11,S13 Closure evidence. Allowable boundary envelope, maximum deltas, required tests, hold points, rollback and new-submission triggers Pass / re-open. Change remains inside validated envelope and required tests pass before release Re-open when: Change outside envelope or failed control/test.
Original source ↗
Source directory

All 21 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Matthew Collins (Quality and Regulatory Executive)Industry · Sep 15, 2026Nathan SabichIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026SichGate Inc.Industry · Aug 22, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026