FDA GenAI discussion / Question 24 of 26

When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?

Full FDA question

For devices built on third-party foundation models, changes to the underlying model may be initiated by the third-party model developer rather than the device manufacturer. How can a manufacturer detect, evaluate, and respond to such changes in a timely manner? What mechanisms—for example, contractual, technical, or through a PCCP—could provide reasonable assurance that third-party developer-initiated changes do not compromise the safety or effectiveness of the device?
Read the FDA discussion paper ↗

30 of 95 submissions reference this question.

All audiences
21 Industry5 Clinicians2 Public / patients2 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/24
Filter by audience
Question 24 · Public feedback

What respondents recommend

15 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Identify and control the model version in use13
Detect supplier updates or unexpected behavior changes12
Retest changed models or provide rollback12
Manage the model as a safety-relevant supplier component1
QRx PartnersIndustry · Aug 24, 2026
Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Sam Rosenthal (Red Kit)

Industry · Sep 9, 2026

Identify and control the model version in use · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 — third-party foundation model changes. The paper frames this as a problem of detecting changes the model developer makes. For an on-device product the answer is technical and complete: the device ships a specific set of weights identified by a cryptographic hash; the app verifies that hash before it will use them; the only way the model changes is a manufacturer-initiated release that goes back through the benchmark. There is no third-party-initiated change to detect. I would ask CDRH to recognise "pinned, hash-verified on-device weights" as a mechanism that resolves Question 24 by construction, and to treat a fine-tuned open-weight model whose base is published under a fixed version as a manufacturer-controlled model for this purpose, not as a live dependency on a third-party service. One closing observation. The paper's framework is built for a world in which the device sits between a patient and a clinician. There is a large population — offshore, underground, at sea, in the backcountry, in every disaster that takes the network down — for whom the device sits between a patient and nothing. I would ask that the eventual guidance say something explicit about that setting, because it is exactly the setting where a careful developer most needs to know what "safe enough" means, and where a rule written for the home-triage case will either be ignored or will keep the product from existing. Thank you for the opportunity to comment. Sam Rosenthal Red Kit
Original source ↗

Brandon Kaplan

Industry · Sep 8, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24: Manage supplier changes and retirement Manufacturers should document the version guarantees, change notices, evaluation opportunities, and retirement arrangements their suppliers provide. A service identifier may cover less than the full configuration that affects behavior; the manufacturer should establish what it covers and which changes a supplier can make without changing it. The manufacturer should make adoption of a replacement an explicit release decision where version selection is available. For services that do not permit deferral, it should demonstrate how its controls manage change-related uncertainty within the response time required for safety. It should restrict the affected capability or choose a different dependency if those controls cannot support the intended use. Behavioral checks can help detect a change. Manufacturers should state their coverage and the failures they might miss, rather than treat a passing sample as proof that a remote service is unchanged. Manufacturers should test a fallback that remains available if the supplier retires the previous version. Alternatives may include another evaluated configuration, restricted functionality, or a human workflow. Human reviewers need the information, availability, and capacity to handle the expected workload. NIST addresses testing fallback arrangements, including manual processing. [3, GV-6.2-006] After an incident, manufacturers should address affected outputs, records, actions, and users as well as restoring software. FDA should assess the adequacy of the chosen controls without requiring self-hosting, access to proprietary model weights, indefinite support for old versions, or one notice period for all devices. 5. Question 20: Supervisory agents Manufacturers should define the supervisor's assigned role, the evidence it can access, and its intervention authority. Retrospective screening, triage for human review, and blocking an action require different evaluations. The manufacturer should demonstrate the protection claimed for the selected role. Brandon Kaplan | Individual capacity Page 4 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Evaluators should use task-appropriate reference criteria with qualified clinical or technical adjudication. They should separate evaluation cases from development and tuning cases, document reviewer independence and conflicts, and define how to resolve disagreement. They should use evidence beyond the primary agent's explanation or the supervisor's agreement to establish correctness. Manufacturers should report missed failures by type, false alarms, response times, and uncertainty. They should test shared failure modes and measure the supervisor's added contribution through a justified comparison, such as matched trials with and without its protection in a safe test environment. Choosing different vendors does not by itself establish independence. Two models receiving the same outdated record could agree while missing the selection error; evaluators should check whether the supervisor can access the evidence needed to detect it. The evaluation should include missed unsafe actions, inappropriate blocking of necessary actions, and human-review workload. Manufacturers should use the sampling safeguards in Section 2 to look for errors in unflagged cases. They should assess changes to the supervisor under the same impact-based process used for the device. Manufacturers should define and test the response to an unavailable, delayed, or inconclusive supervisor. They should prevent an uncontrolled bypass of a supervisor needed for safe operation. The evidence burden should follow the claimed protection; this recommendation does not require human review of every output. 6. Investigation records and documentation FDA's software-submission guidance addresses traceability and risk-management documentation, and its cybersecurity guidance addresses security controls and testing. [2, Section VI.C; 5, Section V] I recommend that manufacturers extend those processes to cover agent execution and reuse relevant evidence. They should connect each critical safety claim to its hazard, control, evaluation, and remaining limitation. Manufacturers should retain records needed to investigate foreseeable failures, including configuration identifiers, relevant inputs and source versions, user-facing outputs, approval events, tool requests and results, and resulting state changes. They should correlate records across the agent and connected systems, with enough timing information to reconstruct the sequence. Investigators should distinguish proposed, attempted, blocked, completed, and uncertain actions, and confirm consequential effects against the affected system's records. Manufacturers should protect those records from unauthorized alteration or deletion, including by the agent. They should detect loss of required records and respond according to its safety impact. An interruption in retrospective logging may permit continued operation under a documented limit; loss of evidence needed for an immediate safety decision may require restricting the affected function. Manufacturers should test both conditions. Manufacturers should justify the patient information retained, its purpose, retention period, and access controls. They may use protected references if authorized investigators can retrieve the relevant historical versions for the required retention period. A current chart alone cannot establish what the device saw earlier. Manufacturers should document gaps and protect any additional content needed for investigation under applicable privacy and recordkeeping obligations. Investigators should rely on observable actions and state changes to establish what occurred. A model-generated explanation presented to a user may itself be relevant evidence, but it does not prove that the described actions occurred. This recommendation does not require disclosure or retention of internal model reasoning. Brandon Kaplan | Individual capacity Page 5 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Manufacturers should rehearse investigations using the records and access arrangements available after deployment. Reviewers should be able to identify the configuration, reconstruct consequential actions, and state unresolved gaps. Manufacturers should preserve the exact content shown to users when its wording matters to safety. Investigators need not reproduce that content by running the model again. Conclusion I recommend that FDA ask manufacturers to support critical safety claims with tested controls and documented limitations. Monitoring and response plans should identify who can intervene, what they can do, and whether they can act within the time justified for the hazard. Manufacturers should support those plans with evidence suited to the device's intended use. Brandon Kaplan References 1. U.S. Food and Drug Administration. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. 2026. See pp. 1, 11-14, 19-23, and 26. Discussion paper, not draft or final guidance. 2. U.S. Food and Drug Administration. Content of Premarket Submissions for Device Software Functions. Final guidance, June 14, 2023. Sections VI.C and VI.J. 3. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, July 2024. MANAGE 4.1; GV-6.2-003 and GV-6.2-006. Cross-sectoral reference. 4. U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final guidance, August 18, 2025. Sections V.D and VI-VIII. 5. U.S. Food and Drug Administration. Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions. Final guidance, February 3, 2026. Section V, especially V.B-V.C. Brandon Kaplan | Individual capacity Page 6 of 6
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Identify and control the model version in use

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 assumes a manufacturer knows which third-party model its device is using. In field measurement conducted on 2026-09-03, that assumption did not hold, and it did not hold on a system built and maintained by an engineer actively auditing it. We submit that CDRH should treat component identification as a prerequisite to change detection, not as a step already completed. • A manufacturer cannot detect a change in a component it has misidentified. Documentation drifts from deployment silently and is not a reliable record of what is running. • The most consequential failure we measured was not in the model. It was upstream of the model, in audio capture, and no model-monitoring program would have seen it. • The failure produced no log entry of any kind. A postmarket monitoring program cannot observe an event the device declines to record. • The residue the failure left behind was indistinguishable from ordinary user behavior. An automated reviewer examining the primary evidence concluded, incorrectly and with confidence, that the user had simply stopped speaking. 1. The prior question: which model is actually running? Before a manufacturer can evaluate a developer-initiated change, it must be able to state which model, at which version, served a given output. We tested that on our own production stack, where the answer should have been trivial. It was wrong. The transcription function in our media pipeline carried an in-code description stating that speech recognition was performed locally by a named open-weights model at a specified file path. The FDA-2026-N-7874 — Question 24 1 The Christman AI Project function body below that description did not call that model. It posted audio to an internal service that forwarded it to an entirely different vendor's hosted speech API. The named local model was present on disk, 465 MB of it, unused by that path. A health endpoint separately reported the local model as ready for a code path that had not used it in weeks. An earlier internal audit of the same codebase had found the reverse error still propagating: nine files across three repositories asserting a vendor arrangement that had already been reversed, and an automated test asserting the false version. This is a first-party system, under active audit, maintained by its author. If component identity cannot be trusted here, it cannot be assumed in a device assembled from vendor SDKs through a procurement chain. Any change-detection obligation built on self-reported component identity inherits this defect. 2. Field measurement: the failure was upstream of the model Three screen recordings were made between 01:27 and 02:42 on 2026-09-03 while dictating into the built-in speech-to-text of a commercial AI assistant application, on a dedicated USB audio interface. The recordings were made because dictation kept terminating mid-sentence. Audio was extracted to 16 kHz mono PCM and examined in contiguous 250 ms windows across the full duration of each file. No sampling was used. The measure is the count of distinct 16-bit sample values per window. A live microphone in an occupied room yields thousands of distinct values per window even during pauses, because room tone is a signal. Measured speech windows in these files carry 7,000 to 12,000 distinct values. A collapse to single digits, with one constant occupying nearly the whole window, is not a quiet room. It is a stream that has stopped delivering samples. Recording Wall time Audio carrying no signal Share of file Held constant REC1 01:27:50 14.00 s of 197.4 s 7.1% 2 REC2 02:09:37 40.25 s of 210.4 s 19.1% 1 REC3 02:41:36 31.00 s of 73.5 s 42.2% 1 The proportion of lost audio roughly doubled every forty minutes over a single session. That is a degradation curve, not a configuration value. The signature at a single transition, REC1 at 171.94 to 172.14 seconds: • 3,186 of 3,200 consecutive samples hold the identical value 2. Seven distinct values across the entire 200 ms window; full range minus 4 to plus 4. • Immediately before, at 170.8 to 171.0 s: 2,948 distinct values, peaks near plus or minus 31,000. Ordinary speech. • Immediately after, at 172.16 s: 3,110 distinct values, minimum minus 32,762, maximum plus 32,767. Full scale, resuming mid-word, with no onset. A speaker does not resume at clipping level with no attack. A stream does. The speaker never stopped; the capture did. This matters for Question 24 because the component that failed was not the generative model and FDA-2026-N-7874 — Question 24 2 The Christman AI Project would not have been examined by any monitoring program scoped to model behavior. 3. The transcriber produced fluent speech from dead input The same audio was transcribed with a widely deployed open-weights speech recognition model (Whisper, small.en, run locally under the submitter's control, word-level segmentation). In REC3, of 136 transcribed words, 39 — twenty-nine percent — fall in windows carrying no live signal. Selected output, with the corresponding measured window: 20.00 - 25.25 s (6 distinct values, 99% constant) → “stopped me right” 29.50 - 31.00 s → “ah now i go back up here and” 44.75 - 48.00 s → “i ended that one real” 61.50 - 66.50 s → “text let’s watch it change” The first entry deserves particular attention. Reassembled, the transcript renders the phrase “it stopped me right there.” The system generated the speaker’s own description of the failure out of the silence that the failure produced. The closing words of the recording were likewise written over nothing. In a second recording, across a dead region spanning 129 to 140 seconds, the transcriber produced a grammatical sixteen-word sentence the speaker did not say. Sample statistics for one second inside that region: eight distinct values across the full second, 89% of samples on a single constant. This behavior falls squarely within benchmarking element S.3 as drafted, which treats presenting information with false confidence as a safety failure. We note that the drafted element addresses confidence in clinical content. We recommend it also address confidence in the premise that input was received at all. 4. Nothing recorded that any of it happened The application maintains a dictation log channel. Its most recent session entries predate these recordings by eight days. For the date in question the file contains a single initialization line and nothing further; its last write preceded the first recording. Three sessions terminated. Between thirty-one and forty seconds of speech lost per recording. Zero log lines. We raise this because Question 24 asks how a manufacturer can respond in a timely manner. A manufacturer cannot respond to an event that generates no record, and a postmarket surveillance obligation that assumes failures are self-reporting will not detect this class at all. 5. An automated reviewer misread the primary evidence, twice We report this because Question 20 asks what considerations apply to the reliability of a machine-based supervisory agent, and because we produced a worked example while preparing this comment. A large language model was used to analyze the recordings above. On first pass it applied an amplitude threshold, observed a crossing, and concluded that the application had ended the session because the user paused. The causation was inverted: the threshold crossing was the stream failing, and the silence was the consequence rather than the cause. Only direct inspection of raw sample values reversed the finding. FDA-2026-N-7874 — Question 24 3 The Christman AI Project On a second occasion the same system read the transcriber’s confabulated sentence, treated it as the speaker’s words, and stated back to the speaker what he had supposedly said. He had not said it. The error was caught only because he was present to deny it. Both errors occurred with the primary file in hand, with no time constraint, and under explicit instruction to verify. Neither would have been caught by a reviewer working from logs or transcripts alone, because both conclusions were consistent with every downstream artifact. We submit that any supervisory agent proposed for postmarket monitoring must be validated specifically against device-failure residue that mimics normal user behavior, and that agreement between an agent and a device log should not be treated as corroboration when both derive from the same failed component. 6. Recommended mechanisms Responsive to the second half of Question 24. 1. A runtime component manifest, asserted rather than documented. The device should emit, per inference, a machine-readable record of which model identifier, version, and endpoint actually served the output, generated by the calling path itself rather than from configuration or documentation. Our own misidentification was possible precisely because the claim lived in prose beside the code instead of being produced by it. 2. Version pinning as a condition of clearance, with unpinned dependency treated as a change. Where a device calls a third-party model that the developer may update without notice, the absence of a pinned version is itself a change-control gap. A PCCP should be permitted to cover a bounded set of pinned versions the manufacturer has re-benchmarked, and should not be available to devices calling an unversioned endpoint. 3. Input-integrity monitoring, specified separately from model monitoring. The failure measured here was in audio capture, upstream of any model. Monitoring scoped to model output would have reported normal operation throughout. For devices ingesting sensor, audio, or image data, we recommend continuous verification that the input stream carries live signal, with loss recorded as an adverse event. 4. Negative evidence must be recorded. Absence of input, premature session termination, and dropped capture should generate log entries with the same obligation as errors. A failure that leaves no trace is unreachable by every postmarket mechanism in Section VI, however well designed. 5. A confabulation-over-null-input benchmark element. Devices generating text from sensor input should be tested against inputs containing no valid signal, with any fluent output treated as a failure rather than scored on plausibility. This is cheap to run, fully automatable, and directly detects the behavior documented in Section 3 above. FDA-2026-N-7874 — Question 24 4 The Christman AI Project 6. Contractual notice is necessary but insufficient. Advance-notice clauses assume the manufacturer can associate a notice with the component actually in the deployed path. Where component identity is not independently verifiable at runtime, notice arrives without a reliable way to determine whether it applies. 7. Evidence and reproducibility The three source recordings, extracted PCM audio, word-level transcripts, frame captures at each failure point, and a full measurement record with window parameters have been retained and are available to CDRH on request. SHA-256 digests of the unaltered source recordings were computed at the time of archiving and are included in that record. Every figure in Sections 2 and 3 is derived from the retained PCM files and can be regenerated from them using the stated window parameters. No figure in this comment requires accepting our characterization of it. We do not offer a word-error rate. An initial comparison between the application’s delivered text and an independent transcript produced a divergence figure that did not survive inspection: most of it proved to be tokenization artifacts of the comparison transcriber rather than alteration by the device. The findings above are anchored to sample-level measurement instead, and do not depend on comparing two fallible transcripts. 8. About this submission The Christman AI Project builds augmentative and alternative communication systems for nonverbal and neurodivergent users, cognitive support for dementia care, and related assistive technology. The submitter is autistic and builds for this population directly. We raise these findings because every protective factor that operated on 2026-09-03 is absent for the users we build for. The failure was noticed because a sighted user was watching the screen at two in the morning, made a recording, and had the means to inspect raw audio. A nonverbal user relying on an AAC device has no screen to watch, no second recording, and no waveform. The sentence the system invents is not an inconvenience to them. It is their voice in the record, and they cannot say that they did not say it. Measurements dated 2026-09-03. Submitted to Docket FDA-2026-N-7874, comment period closing 2026-10-19. Contact: contact@thechristmanaiproject.com FDA-2026-N-7874 — Question 24 5 The Christman AI Project
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24: Third-party model changes The manufacturer must know when a third-party foundation model changes. Contracts should require: Advance notice. Model and version identification. Release information. Safety incident information. A suitable deprecation period. Access to an earlier version. Investigation support. Rollback support. Technical controls should include: Version logs. Fixed production versions where possible. Automated regression tests. Tests before migration. Detection of unexpected behavior changes. Limited initial deployment. Rapid rollback. The manufacturer must assess the complete device after each important model change. When the model provider does not permit sufficient control, the manufacturer has an unresolved device risk. Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices FDA Docket No. FDA-2026-N-7874 Conclusion CDRH should use three assurance layers. Competency must show that the final device can perform its task and remain within scope. Clinical confirmation must show that the device works in the intended clinical setting. Operational assurance must show that the device continues to work after deployment. CDRH should use a clear division between general output and patient-specific output. It should not regulate small differences in wording. Manufacturers should use deterministic controls for boundaries that a system can measure. Postmarket monitoring must connect each signal to an accountable action. The available actions must include investigation, restriction, pause, rollback, and withdrawal. Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
Original source ↗

Krishna Koka

Academia / other · Sep 1, 2026

Identify and control the model version in use

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Couple change control across model, firmware, build-parameter, material, and postprocessing changes, with tiered triggers and pinned model versions. • Require separate surgeon approval and manufacturing release, and a proportionate digital manufacturing record linking configuration to outcome. I appreciate CDRH’s attention to these considerations and would welcome the opportunity to discuss them further. Respectfully submitted, Krishna Sai Koka, M.S., B.S.E.​ Medical Student, New York Medical College (submitted in an individual capacity)​ kkoka@student.nymc.edu Comment on Docket No. FDA-2026-N-7874 — Page 3
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Discussion Question 24 – Changes of Third-Party Foundation Models Third-party foundation models present a particularly significant challenge because safety-relevant changes may occur outside the direct control of the device manufacturer. An assessment of changes in the underlying foundation models should be included explicitly in the postmarket monitoring strategy. This may include specific benchmarking tests to assess the impact of changes in the foundation model on resulting clinical performance. Additionally, these steps should address effects on the interaction between the GenAI-enabled device and users as well as the deployment environment. This should be aligned with the competency-based approach already discussed. Major changes in foundation models or other components may necessitate a downgrading of the currently assumed competency level. The GenAI-enabled device may need to return to a stage with greater supervision. This would amount to a requalification period in a more controlled environment. A more independent use of the GenAI system can be re-established when defined evidence gates are successfully passed. An additional requirement should be case-level reconstructability. For meaningful investigation of complaints, adverse events, or performance signals, the manufacturer should be able to reconstruct, to the extent technically feasible, the relevant device configuration at the time of use. This may include the foundation-model version, prompts, guardrails, retrieval sources or knowledge-base version, tools, orchestration logic, relevant runtime parameters, and applicable clinical-environment information. Exact reproduction of every stochastic output may not always be feasible. Nevertheless, the historical system configuration should remain sufficiently traceable to permit reliable investigation and reassessment of safety-critical behavior. Concluding remarks to Section VI Overall, postmarket monitoring for GenAI-enabled devices should be understood as an active lifecycle control system rather than a periodic performance check. Its purpose should be to maintain assurance of the product- specific benefit-risk profile, verify continued effectiveness of the safeguards underlying the device's criticality, detect changes in both the device and its clinical environment, and trigger proportionate corrective action. Importantly, the rate of permitted technological and operational change should remain compatible with the ability of the monitoring and regulatory system to detect, understand, and control its consequences. Section VII – Other Topics
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 - Third-party foundation-model changes Manufacturers need technical and contractual mechanisms that make third-party model change observable. Useful mechanisms include model/version attestation, cryptographically or otherwise reliably verifiable version identifiers, update notifications, structured change manifests, behavioral release notes, compatibility testing, regression suites, and contractual access to safety-relevant change information. Where technically and contractually feasible, sponsors should consider version pinning, qualification windows, or equivalent change-detection and release controls that prevent uncharacterized model changes from entering high-consequence clinical use without detection and evaluation. If the model provider changes the underlying system unexpectedly, the device should have a defined response, such as entering a restricted mode, suspending affected functions, or requiring requalification before continued high- consequence use. An unchanged API endpoint or product name should not be treated as proof that the clinically relevant model behavior is unchanged. SECTION VII - OTHER TOPICS
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 — Changes to third-party foundation models “The foundation model changed” cannot be a liability shield. A manufacturer cannot credibly guarantee safety if the model underneath its device can change without notice. FDA should expect contractual and technical guardrails — version identification, advance change notification, rollback capability, and audit logs — so a manufacturer marketing a device built on someone else's model still owns that device's clinical behavior.
Original source ↗

QRx Partners

Industry · Aug 24, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback · Manage the model as a safety-relevant supplier component

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24: Third-Party Foundation Models and Change Control The challenge presented by third-party foundation models is significant, but the underlying regulatory problem is not entirely novel. Medical device manufacturers already incorporate software components they did not develop and over which they may have limited development information or control. FDA should consider building on established principles for managing Software of Unknown Provenance (SOUP), including IEC 62304 and existing FDA software lifecycle expectations. Manufacturers should evaluate the third-party component within the finished device, identify how reasonably foreseeable failures or changes could contribute to hazardous situations, evaluate resulting risks, and establish appropriate controls. The objective should not be complete visibility into or control over the foundation model, but sufficient information and controls to identify, evaluate, and respond to changes that could affect device safety or effectiveness. Controls could include supplier change notification, model version control, regression benchmarking, behavioral monitoring, guardrails, and other device-level controls. Foundation models introduce an additional challenge because a remotely hosted model may change without an identifiable modification to the manufacturer's software. Mechanisms may therefore be needed to detect significant behavioral or performance changes when supplier notification is unavailable. Consistent with FDA's emphasis on evaluating the final user-facing device rather than the foundation model in isolation, expectations should focus on the manufacturer's ability to establish and maintain reasonable assurance of safety and effectiveness of the finished device. Overall Consideration GenAI creates new failure mechanisms and assurance challenges, but these do not necessarily require replacing established risk-based medical device principles. New approaches should remain focused on intended use, reasonably foreseeable failure, residual risk, and clinically meaningful evidence rather than attempting to evaluate every capability or possible behavior of an underlying generative model.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 — What happens when the underlying third-party foundation model changes? Response The organization placing an AI-enabled healthcare product into use cannot disclaim responsibility because a third-party model provider changed the underlying technology. If a manufacturer or deployer elects to build upon an external foundation model, management of that dependency is part of the safety obligation. At minimum, there should be mechanisms addressing:  notification of consequential model changes;  version identification;  regression testing;  revalidation of safety-critical functions;  monitoring after updates;  documentation of which version produced which output;  rollback capability where feasible;  contractual requirements concerning change notification;  and contingency planning if an underlying model is materially altered or withdrawn. The patient cannot reasonably be expected to understand that yesterday’s healthcare AI and today’s healthcare AI may share a brand name and interface while relying on meaningfully different underlying model behavior. That is an enterprise governance issue. Not a patient responsibility. Again: Contracting out a technological component does not contract away responsibility for the healthcare product built upon it. VII. Agentic AI and the Clinical-Adjacent Safety Gap
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 24 — Third-party model changes initiated by the model developer In my assessment this is the most significant unaddressed gap in the current regulatory picture, and I am glad the paper names it. My understanding of prevailing commercial practice — which I would encourage CDRH to verify directly with model providers, as it varies by provider and is changing quickly — is that hosted model endpoints are frequently updated without individualized customer notice; that version pinning is offered inconsistently and often with limited-duration guarantees; and that pinned versions are deprecated on provider-determined timelines that may be shorter than a device lifecycle. If that characterization is broadly accurate, then a device manufacturer relying on a hosted third-party endpoint may have its device’s behavior changed without its knowledge and without any mechanism to detect the change, which is not a condition any quality system can accommodate. I recommend: 15 of 19 Docket No. FDA-2026-N-7874 Treat version pinning as a design control expectation. CDRH should state that a GenAI-enabled device relying on a third-party model is expected to execute against a pinned, version-identified model endpoint, and that the inability to pin — or reliance on a provider that does not offer pinning with change notice — is a design deficiency to be addressed, not a residual risk to be accepted. This is the single highest-leverage statement CDRH could make in this area, and it would immediately reshape provider offerings in the medical device segment. Require contractual change notice with a defined minimum period, sufficient to complete the Tier 2 re- benchmarking described above before a change takes effect, together with a minimum deprecation window for pinned versions. Require continuous canary monitoring, because contracts cannot be verified in real time. A small, fixed probe suite — a set of inputs with known expected outputs — executed against the live endpoint at high frequency provides direct detection of silent upstream change. This is inexpensive, technically straightforward, and the only mechanism I am aware of that verifies rather than assumes endpoint stability. I recommend CDRH describe it as an expected element of postmarket monitoring for any device on a third- party hosted model. Treat provider-forced migration as a Tier 3 change. When a provider deprecates a pinned version and the manufacturer must migrate, the resulting device is running on a different model. That it was involuntary does not change what it is, and the paper should foreclose the argument that forced migrations warrant lighter treatment. Manufacturers should be expected to plan for this contingency — including maintaining a validated fallback — as part of design planning rather than handling it as an emergency. I would also note that this is where a Foundation Model MAF (Question 25) would deliver the most value: a standing, current record of a model’s versioning, deprecation, and change-notification practices would let CDRH and manufacturers assess dependency risk directly rather than by inference. IV. Other topics (Section VII; Questions 25–26)
Original source ↗

Supernova Technologies

Industry · Aug 18, 2026

Detect supplier updates or unexpected behavior changes

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Point Three: Independent Detection of Performance Degradation (Questions 19 and 24) The paper names performance degradation over time as one of the core risks of GenAI enabled devices, and proposes performance degradation monitoring as a postmarket approach. I recommend that CDRH require degradation detection mechanisms to operate independently of the deployed device's own output or reporting pathway. A device that is degrading, whether due to a change in the underlying foundation model, drift in the input population, or another cause, may not reliably signal its own decline through the same channel used to generate its outputs, particularly for GenAI enabled devices where a single failure mode can affect both the device's clinical output and its own self assessment of that output. This concern is closely related to Question 24, regarding changes initiated by a third party foundation model developer rather than the device manufacturer. In both cases, the manufacturer's ability to detect a problem depends on a signal that is generated or mediated by the same system that may be experiencing the problem. I recommend that degradation monitoring specifications require at least one detection mechanism, such as periodic re benchmarking against an independent reference set or sampled clinician review as described in Section VI.A, that does not depend on the device's own reporting of its performance. Conclusion The concerns raised above are not arguments against the competency based framework or against reliance on postmarket monitoring generally. Postmarket monitoring, done well, is likely necessary given the practical limits of premarket testing for open ended systems described elsewhere in this paper. My recommendation is narrower: that CDRH require monitoring and disclosure mechanisms to demonstrate their own reliability and completeness as a condition of being credited toward reduced premarket evidence, rather than treating the existence of a monitoring program or a voluntary disclosure mechanism as sufficient on its own. Given this year's demonstrated pattern of AI developers failing to detect problems in their own systems despite strong incentives to do so, this distinction is likely to matter in practice. Thank you for the opportunity to comment.
Original source ↗

Richard Pescatore, DO (BellyMD)

Industry · Aug 18, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Questions 24 and 25: third-party foundation models. A small manufacturer cannot compel a foundation model developer to disclose changes or give advance notice, and contractual leverage is concentrated in the largest sponsors. Four mechanisms would help. Treat model version pinning as a baseline design expectation, with developers disclosing deprecation timelines. Advance the voluntary Foundation Model MAF program with update notification commitments and healthcare-relevant evaluation summaries as core content, and signal that devices built on MAF-holding models will see more predictable review; that signal creates the developer's incentive to participate. Endorse sponsor-side re-benchmarking gates, under which no underlying model change enters production until the premarket benchmark battery has been re-run and passed, consistent with Section VI.C. And define a PCCP category for like-for-like model version upgrades validated through that prespecified re-benchmarking. Together these convert an uncontrollable third-party dependency into a controlled, auditable change process inside the sponsor's quality system.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 24 - Third-party foundation-model changes Trace ID. TR-Q24 | FDA Q24; Sec. VI.D; App. B; pp. 21-22 / 29-30 BCR response. Use contractual notification plus technical detection: version identifiers/fingerprints, canary probes, shadow/delta tests, API behavior checks, change quarantine, rollback, and requalification. If the sponsor cannot detect a safety-relevant dependency change, the branch is not closed. BCR rule basis. BCR-R01,R03,R11,R13,R16 Solution-stack link. S10,S11 Closure evidence. Provider notice + fingerprints + canary/shadow delta tests + quarantine/rollback Pass / re-open. No unrecognized safety-relevant dependency change reaches production Re-open when: Any provider/model/API/safety-behavior change.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Identify and control the model version in use · Detect supplier updates or unexpected behavior changes · Retest changed models or provide rollback

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 24: Third-Party Foundation Model Changes Manufacturers bear responsibility for the performance of their devices, but they cannot fulfill that responsibility if they lack notice of changes to the foundation models on which their devices are built. FDA should require — as a condition of market authorization for devices built on third- party foundation models — that manufacturers demonstrate contractual notification rights providing a minimum of ninety days advance notice of any material model version change, including version deprecations. Shorter notice windows are inadequate for conducting validation studies, executing change control procedures, and notifying FDA. Technically, FDA should require manufacturers to implement model versioning pins with pre- production validation gates: no foundation model update should reach a deployed medical device without first clearing a defined validation protocol. Automated deployment of foundation model updates to live medical devices — without validation — should be explicitly prohibited. For devices where contractual notification rights cannot be secured from the foundation model developer, FDA should impose correspondingly higher premarket evidence requirements to account for the elevated ongoing risk. VI. SECTION VII: FOUNDATION MODEL MAFs AND AGENTIC AI — RESPONSES TO DISCUSSION
Original source ↗
Source directory

All 30 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Brandon KaplanIndustry · Sep 8, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Matthew Collins (Quality and Regulatory Executive)Industry · Sep 15, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026QRx PartnersIndustry · Aug 24, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Supernova TechnologiesIndustry · Aug 18, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Sihem KhelifaClinicians · Sep 9, 2026Qiong LiuPublic / patients · Sep 11, 2026Saurabh SharmaPublic / patients · Sep 14, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026