Allow it only under defined conditions
Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.
Read the source passage
The question as posed CDRH is considering whether it may be appropriate to accept greater premarket uncertainty regarding a GenAI-enabled device's benefit-risk profile through greater reliance on postmarket monitoring. Under what conditions might such an approach be appropriate, and what characteristics of a monitoring program would need to be in place to justify reduced premarket evidence? Are there device types or risk profiles for which this approach would not be appropriate? Summary of position We support the trade in principle. A total product lifecycle approach is the right instinct for devices whose behavior is open-ended and whose models change after clearance. Our comment is about the condition on which the trade depends, and about one class of device for which we recommend it not be available at all. • A reduction in premarket evidence is only a trade if the monitoring program can detect the failure the premarket evidence would have caught. Where it cannot, nothing has been exchanged — certainty has been given up for coverage that does not exist. We recommend the reduction be granted against demonstrated detection capability, never against the existence of a monitoring program. • We have measured a failure mode that a periodic monitoring program cannot see by construction. It occurs within a single session and resets at the session boundary, so every scheduled reassessment measures the healthy state. A program built on periodic re-benchmarking should not be able to fund a premarket reduction for any failure mode of that shape. • Monitoring that reads the device's own account of its work inherits the device's errors. If premarket evidence is reduced on the strength of postmarket monitoring, and that monitoring is a self-report, the reduction rests on the device vouching for itself. • We recommend the approach be unavailable where the intended user cannot self-verify or self-report, where the device operates without connectivity, and where the harm completes inside the detection interval. Each Docket FDA-2026-N-7874 · The Christman AI Project and Robotics Division Page 1 of these describes the assistive and communication devices we build, and each removes a safeguard the trade quietly assumes is present. 1. The unstated premise in the trade Accepting greater premarket uncertainty in exchange for postmarket monitoring is a sound instrument, and it is already how much of the device program works. But the exchange carries a premise that is rarely written down: that the residual uncertainty is of a kind postmarket monitoring can resolve. For most device attributes that premise holds. A durability question, a failure rate, a rare adverse event — these accumulate in the field and become visible with time and volume. Postmarket surveillance is the correct instrument because the signal is statistical and time reveals it. The failure modes we have measured in generative systems are not of that kind, and time does not reveal them. They are individually invisible, they do not aggregate into a rate, and they are erased by the very act of reassessment. We set out the measurement below, because the recommendation follows from it rather than from a position. 2. The measurement, and why it defeats a periodic program On 2026-09-03 we made three recordings across seventy-four minutes of continuous work on one unchanged audio interface, with no configuration change between them. Input carrying no live signal accounted for 7.1 percent of the first file, 19.1 percent of the second, and 42.2 percent of the third. The proportion of lost input roughly doubled between each recording. The consequence for Question 18 is structural rather than a matter of rigor. A re-benchmark is a fresh session. A reassessment run at any point on any day would have opened a new session and measured something close to the 7.1 percent state. The degradation is a function of elapsed time within a session, and a monitoring cadence measured in weeks or quarters cannot observe a quantity that resets in minutes. A separate recorded session on 2026-09-04, eleven minutes twenty-four seconds, showed the second half of the problem. The device produced false statements in three separate turns while its presentation improved steadily across the session — by the ninth minute it was citing governing rules by name, disclosing source ages, and correcting itself unprompted. A sample drawn late in that trajectory scores the device as more disciplined than a sample drawn early, while the error rate is unchanged. A monitoring program that samples will report the better number, and will report it in good faith. We state the boundary on this evidence in Section 6. It establishes that these failure modes occur and can be measured. It establishes nothing about how often, and we offer no rate. 3. Conditions we recommend attaching to any reduction Responsive to the first and second parts of Question 18. We propose three conditions, each stated so that a reviewer can deOriginal source ↗