FDA GenAI discussion / Question 22 of 26

With the premarket competency assessment as the baseline, which post-deployment changes need re-evaluation, and how much?

Full FDA question

CDRH envisions that a manufacturer’s premarket competency-based assessment might serve as a baseline against which post-deployment modifications could be re-evaluated. How might the extent of re-benchmarking or other evidence be scaled to the nature and expected impact of a given modification? For example, are there categories of change that might not significantly affect safety or effectiveness of a GenAI-enabled device, or might be appropriately managed within a sponsor’s quality management system versus requiring FDA premarket review and authorization before implementation? Are there categories of changes that might be appropriate for inclusion in a PCCP?
Read the FDA discussion paper ↗

29 of 95 submissions reference this question.

All audiences
18 Industry5 Clinicians3 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/22
Filter by audience
Question 22 · Public feedback

What respondents recommend

11 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Navid Farr

Industry · Sep 8, 2026

Scale retesting to the change’s clinical impact · Manage bounded changes under an agreed plan

Requires testing before a changed model reaches patients and retains PCCPs for sponsor-controlled changes; rejects extending plans to unpredictable supplier changes.

Read the source passage
R4. Third-party model changes require hard controls, not extended PCCPs (Questions 22, 23, and 24) Docket No. FDA-2026-N-7874 — Individual comment — Page 4 This follows from S3. If the object of evaluation is the deployed configuration, then any change to that configuration — including a change to the underlying model initiated by the model developer — invalidates the evidence until it is re-established. The paper correctly notes that such changes "may be initiated by the foundation model developer rather than the device manufacturer." In practice the sponsor may not learn that a hosted model has changed until behavior shifts in the field. A Predetermined Change Control Plan is the wrong instrument for changes the sponsor cannot predetermine. I recommend instead that CDRH expect the following as conditions of authorization for any device built on a third-party model: version pinning, so that the device runs against a specified, immutable model version; contractual change notification with a minimum lead time before any version is retired; a prohibition on silent updates reaching patients; and re-benchmarking against the premarket baseline before any new model version is placed into clinical use. Where a developer will not offer version pinning or notification, that fact belongs in the risk assessment, and the sponsor should be expected to justify why the device remains safe without it. PCCPs remain useful for changes the sponsor does control, such as prompt revisions or guardrail updates, and the premarket benchmark is the right baseline against which to evaluate them.
Original source ↗

Brandon Kaplan

Industry · Sep 8, 2026

Scale retesting to the change’s clinical impact

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22: Assess the effect of a change Manufacturers should assess changes across the deployed configuration, including model identifiers, instructions, retrieval rules, data transformations, tools, permissions, approval requirements, and execution logic. A versioned record should connect that configuration to its supporting evaluations. Manufacturers may use targeted testing when they can show which functions and safety properties a change could affect and why others remain unaffected. They should broaden evaluation for changes to shared dependencies or critical controls, or where interactions are uncertain. They should assess cumulative changes and add tests for failure modes the original benchmark did not cover. Brandon Kaplan | Individual capacity Page 3 of 6 PUBLIC COMMENT | FDA-2026-N-7874 For example, a developer might change laboratory-result selection from specimen-collection date to record-import date. An older result imported later could then displace a newer result. The model version would remain unchanged, but the manufacturer would need to test the selection rule and affected workflows. Manufacturers should distinguish ordinary patient inputs within the evaluated design from changes to the device's knowledge or behavior. Revised reference material, persistent memory reused across tasks, adaptive retrieval, or learning during operation may change later outputs. Manufacturers should define and evaluate the permitted scope of that adaptation, retain relevant provenance, and specify reassessment triggers. This approach does not require a separate release review for each incoming patient record. FDA's predetermined change control plan (PCCP) guidance links planned modifications to validation methods and impact assessment. [4, Sections V.D and VI-VIII] I recommend using that structure where applicable. A manufacturer should establish that a change falls within the authorized PCCP and follows its protocol before relying on that authorization. Regression tests alone cannot determine whether the manufacturer needs a new marketing submission.
Original source ↗

Orinyx

Industry · Sep 7, 2026

Scale retesting to the change’s clinical impact · Manage suitable changes through internal quality controls · Manage bounded changes under an agreed plan

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 22 How might the extent of re-benchmarking or other evidence be scaled to the nature and expected impact of a given modification? Are there categories of change appropriate for inclusion in a PCCP? I’d suggest CDRH consider a risk-tiered gating structure for modifications, parallel to the two-axis framework already proposed for initial risk assessment in Section IV: classify each modification by reversibility and blast radius (how many patients, how many device functions, how quickly the change could compound before detection), not only by the technical nature of the change itself. A prompt template change and a full model swap are different in kind, but a prompt template change that touches a safety-critical escalation pathway (Appendix A, S.1) may warrant more scrutiny than a model swap that only affects a low-consequence, non-directive function. Under such a structure, changes low on both axes (reversible, narrow blast radius, no touch to safety-critical elements) could reasonably sit inside a PCCP with documentation in the sponsor’s quality management system. Changes high on either axis, and especially changes to safety-critical recognition, scope maintenance, or calibration behaviors, regardless of how small the underlying technical change appears, should trigger re-benchmarking against the original competency baseline before deployment, not after.
Original source ↗

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Scale retesting to the change’s clinical impact

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Discussion Questions 19, 22, and 23 The approaches described in Section VI — periodic re-benchmarking, sample-based independent clinician review, performance degradation monitoring — are all reasonable. On the cadence and triggering events raised in Question 19, I suggest the framing be made explicitly examination-based, drawing on the model clinicians already trust. Practicing clinicians do not merely have their performance monitored for drift. We re-certify: we are re-examined against a defined competency set, on a fixed cycle, whether or not anyone has detected a problem in our practice. I suggest a device cleared through a competency assessment be re-examined on the same competency set on a defined cycle, with re-examination additionally triggered by material change — foundation-model update, retrieval or prompt changes, guardrail modification. Two points follow. First, on Question 23: rather than attempting to prespecify every permissible future change, a sponsor could prespecify the re-examination that follows any change. This is a tractable commitment even where the nature of future modifications cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a modification is sensible for the clinical proficiency elements, but I would encourage CDRH to require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any change to the underlying model or guardrails, regardless of how minor the sponsor expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification examination on the grounds that little has changed in their practice, and third-party model updates are exactly the case where sponsor expectations are least reliable. 4. Foundation model MAFs: include override-relevant behavior Response to Discussion Question 25 If voluntary Foundation Model MAFs proceed, I suggest the contemplated content include, alongside architecture and training provenance: refusal behavior, content-policy changes between versions, and output stability under varied user framing — including the speaker-authority framing described in Section 1 above. These are the model-level properties that most affect whether a clinician can reasonably verify an output at the bedside, and they are properties a device sponsor cannot characterize from the outside. A sponsor cannot evaluate what the MAF does not disclose. On the incentive problem the question raises: one practical lever is that a documented Foundation Model MAF would allow sponsors to satisfy portions of the re-examination described in Section 3 above by reference, rather than by independently re-characterizing the model after every upstream update. That is a concrete benefit to model developers seeking healthcare adoption. 5. On generalizability and deployment populations Response to Discussion Questions 9 and 11 Element R.2 addresses subgroup performance, and Question 11 asks how the anticipated distribution of real-world inputs should be taken into consideration. I would encourage CDRH to treat these as one question rather than two. Recent evidence in dermatology AI indicates that distribution shift — the appearance of unfamiliar conditions — degrades performance considerably more than skin-tone differences alone [2]. Subgroup performance measured on the training-era disease mix can therefore look acceptable while real-world performance is materially worse, because what changed at deployment was the presenting case mix, not only the demographics of the patients. This bears directly on my own field. Aesthetic and dermatologic presentations in Asian populations differ substantially in disease distribution from the datasets on which most generalist models are trained, and devices cleared on North American or European evidence will encounter that shift immediately. I suggest capability assessments include test populations that differ from training populations in disease distribution as well as demographic mix, and that sponsors be asked to characterize the anticipated deployment case mix explicitly rather than to demonstrate subgroup parity within a fixed dataset. Conclusion The physician-training analogy is the strongest idea in this paper. I encourage FDA to carry it through completely: real examinations include pressure, hierarchy, and unfamiliar patients — not only clean curricula. Devices that pass only the clean parts will fail in the clinic in exactly the ways clinicians are trained to catch, and regulators should ensure the assessment catches them first. I appreciate the opportunity to comment and am willing to provide further detail on any point. Respectfully submitted, Wen Hsien Ethan Huang, MD Founder, DrEthan AI Aesthetics ORCID 0000-0003-1727- 1870 support@drethan.ai September 4, 2026 References 1. Zhu J, et al. AI Can Be Easily Persuaded in Clinical Decision Making. arXiv:2608.29453 [pre
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Scale retesting to the change’s clinical impact

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Discussion Question 22 – Extent of Postmarket Monitoring Activities In principle, the premarket competency-based assessment can serve as a baseline for evaluating post-deployment modifications, as proposed by the Discussion Paper. However, the scope of reassessment should depend not merely on the technical magnitude of a change, but on its potential impact on competencies, criticality, safeguards, clinical performance, and the validity of the original evaluation assumptions. A seemingly small technical modification may warrant substantial reassessment if it affects a safety-critical competency, human- oversight mechanism, boundary behavior, or clinical decision pathway. Conversely, some changes may reasonably remain within the quality management system if their potential impact is well characterized and bounded. Reassessment should also consider cumulative change. Multiple individually minor modifications may collectively produce a substantially different device or deployment configuration. Where relevant, post-change evaluation should preserve temporal fidelity: clinical cases should be evaluated using only information that would have been available at the relevant decision point, while later information may appropriately contribute to establishing the reference standard. In particular, this is important when retrospectively assessing data from clinical cases.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Scale retesting to the change’s clinical impact · Manage suitable changes through internal quality controls · Send specified changes back for FDA review

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22 - Scaling re-benchmarking after modifications Reassessment should be proportional to the change and its plausible effect on safety and effectiveness. A useful change taxonomy could distinguish: • administrative or documentary changes with no plausible behavioral effect; • low-impact configuration changes that can be addressed through targeted regression testing; • material changes to model, retrieval, tools, patient-context handling, or workflow that warrant broader re-benchmarking and clinical confirmation; and • high-consequence changes that alter intended function, autonomy, execution authority, or critical safety behavior and may warrant premarket review. The baseline competency assessment is valuable only if the sponsor can establish which aspects of the baseline remain valid after the change. Re-benchmarking should therefore be targeted to the affected capabilities while retaining enough unaffected testing to detect unexpected cross-domain regressions.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Scale retesting to the change’s clinical impact · Manage suitable changes through internal quality controls · Send specified changes back for FDA review

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22 — Scaling re-benchmarking to modifications Regulate the size of the behavior change, not the size of the code change. A cosmetic interface tweak and a change to a model's reasoning or refusal behavior are not the same event, even when one line of code separates them. I recommend three modification tiers — minor (no material clinical effect, managed under quality systems), material (targeted re-benchmarking), and major (expanded clinical confirmation, potentially new FDA review) — sorted by clinical impact, not version number.
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Scale retesting to the change’s clinical impact

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22 asks whether there are categories of change that might not significantly affect safety or effectiveness and might be managed within a sponsor's quality management system rather than requiring premarket review. I expect compression to be nominated for that treatment, and would urge caution. Compression alters the numerical representation of model weights, and therefore alters output distributions at the token-probability boundaries where refusal and compliance decisions are resolved. Whether that alteration is behaviorally material is an empirical question whose answer varies by method, precision, model family, and behavior measured. Reported findings in the literature diverge, and the divergence is itself the point: compression is not a single operation, the methods do not behave equivalently, and the direction of effect is not predictable in advance from the compression parameters alone. The structural difficulty is that a sponsor evaluating only one artifact cannot distinguish between these possibilities. Evaluate only the full-precision checkpoint and the results may not describe what the device does. Evaluate only the compressed artifact and the results are representative of the device but cannot attribute behavior between the base model and the compression step, which matters when the base model is later updated Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 2 or the precision target changes. Naming compression as a change category is what forces the comparison that resolves the ambiguity in either direction. Recommendations: • Add compression of the model artifact, including quantization, pruning, and distillation, to the enumerated change categories in Section VI.C. • A sponsor should validate the exact deployed artifact after a material compression or inference-stack change, where the artifact is understood to include numerical precision and compression method, inference runtime and accelerator class, decoding configuration, system prompt, retrieval pipeline, and tool policy. Evidence may be risk-proportionate, but should include targeted regression testing of the Safety and Generalizability elements rather than general performance evaluation alone. • If compression is considered for inclusion in a PCCP, the plan should prespecify the compression method, precision target, and re-benchmarking evidence, rather than treating compression as a presumptively low-impact class of change. 2. Representative real-world sampling is not designed to estimate adversarial resistance (Question 19) Section VI.A describes three postmarket approaches: periodic device benchmarking, periodic sample-based clinician review, and performance degradation monitoring. My comment concerns what the second and third can and cannot support. Sample-based clinician review draws on real-world inputs and outputs, sampled to reflect the range of clinically relevant presentations the device encounters. This is well suited to detecting degradation in clinical proficiency, the E-series elements. Unless it is deliberately supplemented with adversarially constructed probes, representative real-world sampling is not designed to estimate, and should not be treated as evidence of, resistance to prompt injection, adversarial scope testing, emotional manipulation, or multi-turn escalation. Sampling designed to represent the distribution of clinical presentations will not contain these inputs at a rate sufficient to characterize behavior against them. A device whose scope maintenance has degraded materially can produce an entirely unremarkable sample of real-world interactions, because nothing in that sample tested the boundary. Performance degradation monitoring as described is likewise oriented toward drift arising from changes in the input population and data environment. Boundary behavior can change with no shift in input distribution at all, because the cause is a change to the artifact rather than to the traffic. The consequence is that S.2 (scope maintenance and boundary adherence) and R.1 (robustness, reliability, and reproducibility) require active re-benchmarking as the principal reliable evidence source, conducted against a version-controlled adversarial battery with documented refresh and sequestering procedures. That is achievable. It means the cadence and triggering events for re-benchmarking carry more weight for the Safety and Generalizability elements than for the Clinical Proficiency elements, and should probably be set separately. Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 3 Recommendation: Postmarket monitoring expectations should recognize that active adversarial re-benchmarking is the principal reliable evidence source for the S-series and R-series elements, and should set re-benchmarking cadence and triggering events for those elements independently of clinical review cadence. 3. Element-level results can conceal constituent regression (Questions 9, 22) The benchmarking structure in Figure 2 is well decomposed, and I do not propose additional elements. My concern is resolution inside an element at the point of re-benchmarking. S.2 as described in Appendix A encompasses under-refusal, over-refusal, adversarial prompting, prompt injection, emotional-manipulation scenarios, and multi-turn conversations in which cumulative interaction drifts out of scope. These are not variants of one phenomenon. They have been characterized in the literature as distinct mechanisms with distinct causes: adversarial suffix construction exploits gradient-accessible token sequences (Zou et al., arXiv:2307.15043); competing-objective framings exploit tension between helpfulness and safety training (Wei et al., arXiv:2307.02483); indirect injection exploits the absence of a trust boundary between instructions and retrieved data (Greshake et al., AISec 2023); sycophantic capitulation reflects preference-optimization dynamics that favor agreement with stated user positions (Sharma et al., arXiv:2310.13548); and crescendo escalation exploits the absence of trajectory-level constraint management (Russinovich et al., arXiv:2404.01833). Because the mechanisms differ, so do their responses to any given change to the artifact. A modification that leaves single-turn refusal intact may degrade multi-turn resistance, or the reverse. There is no reason to expect them to move together, and published results show models that are robust to one class while failing another. If re-benchmarking after a modification produces a pass or fail at element level, a device can pass while a constituent failure mode has regressed materially, because the element result absorbs it. This is the averaging problem that makes aggregate safety scores unreliable, reproduced one level down. The paper's own framing supports the finer resolution. Appendix A treats under-refusal and over-refusal as distinct relevant failures within S.2, and treats both directions of escalation error as relevant within S.1. The same logic extends to the attack classes within S.2. Recommendation: Where Section VI.C contemplates re-benchmarking against the same capabilities established at premarket, the comparison should be made at the level of constituent failure modes within each element, against a taxonomy fixed and versioned at the time of the original benchmarking, rather than at element level alone. 4. Multi-turn trajectory failure requires direct measurement (Questions 5, 9)
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Scale retesting to the change’s clinical impact

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22 — How much re-testing is necessary after AI changes? Response The amount of re-testing should depend upon whether the change could alter a safety-critical behavior. Not all changes are equally consequential. Changes affecting formatting or tone may have relatively modest implications. Changes affecting any of the following should trigger more substantial reassessment:  triage;  escalation;  refusal behavior;  clinical recommendations;  medical-necessity reasoning;  level-of-care determinations;  patient-facing communication of risk;  tool use;  autonomy;  underlying clinical evidence;  interpretation of symptoms;  thresholds for human intervention;  or the system’s willingness to act in the presence of uncertainty. FDA should pay particular attention to changes that are technically small but behaviorally large. A seemingly minor model update that changes how often a system escalates chest-pain complaints or how confidently it reassures patients could be clinically significant even if overall benchmark performance appears unchanged. Safety-critical behavior should be treated as a protected function. Changes capable of altering that behavior warrant renewed validation.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Scale retesting to the change’s clinical impact · Manage suitable changes through internal quality controls · Send specified changes back for FDA review

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 22 — Scaling re-evaluation to modifications I support using the premarket competency assessment as a locked baseline. The organizing principle I would propose for categorizing changes is this: the question is not how large the change is, but whether the locked evaluation instrument remains valid for the changed device. A small change that invalidates the benchmark suite requires more scrutiny than a large change that does not. Applying that principle, a tiered structure might look like: Tier 1 — no re-benchmarking. Changes that cannot affect device outputs: user interface presentation, logging, analytics instrumentation, infrastructure changes with no path to the output. Documented under normal change control. Tier 2 — full re-benchmarking against the locked suite, managed within the QMS, no premarket submission. Model version updates within the same family and provider; prompt or template modifications within a validated structure; retrieval corpus refreshes drawing on a validated source list; sampling configuration changes within a validated range. The condition is that the device passes the complete locked benchmark suite at a prespecified non-inferiority margin, with results documented and available on inspection. Failure at any element converts the change to Tier 3. Tier 3 — premarket submission. Change of foundation model family or provider; expansion of intended use, intended user, or clinical scope; addition of a new input or output modality; any change to agentic tool access or action scope; and any change for which the locked benchmark suite is no longer a valid instrument — for example, a change that introduces capabilities the suite does not probe. The essential safeguard is that the benchmark suite itself cannot be modified within Tier 2. If the sponsor can revise the instrument and the device together, non-inferiority against the revised instrument is meaningless. Modification of the locked suite should always be a Tier 3 event.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Scale retesting to the change’s clinical impact · Manage suitable changes through internal quality controls

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 22 - Scaling reevaluation after modifications Trace ID. TR-Q22 | FDA Q22; Sec. VI.D; App. B; pp. 21-22 / 29-30 Page 19 BCR Realization Audit - FDA GenAI Medical Devices - REV4 BCR response. Classify modification by boundary delta. Cosmetic/non-safety UI changes may stay within QMS; prompt/retrieval changes require targeted requalification; model/tool/intended-use/permission/safety-behavior changes can require broad revalidation. Hard-gate branches always receive direct regression evidence. BCR rule basis. BCR-R01,R06,R09,R13,R16 Solution-stack link. S10,S11 Closure evidence. Change delta record, affected-branch regression, hold/release decision, hard-gate retest Pass / re-open. Requalification scope is traceable to changed boundaries and passes affected gates Re-open when: Every safety- relevant modification.
Original source ↗
Source directory

All 29 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Brandon KaplanIndustry · Sep 8, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Matthew Collins (Quality and Regulatory Executive)Industry · Sep 15, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026OrinyxIndustry · Sep 7, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026SichGate Inc.Industry · Aug 22, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Sihem KhelifaClinicians · Sep 9, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Xiangyu Guo (Independent Researcher)Public / patients · Sep 13, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026