Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals
Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.
Read the source passage
The question as posed Please comment on the potential approaches to postmarket performance evaluation, including periodic re- benchmarking, sample-based clinician review, and performance degradation monitoring. What additional approaches should CDRH consider, and how should the cadence and triggering events for reassessment be determined? Summary of position We support all three approaches and we do not think any of them, on a periodic cadence, can detect the degradation we have measured. Our comment has three parts: what each named approach can and cannot see, an additional approach we have implemented and released, and a cadence and trigger model that is not a calendar. · Every instance of degradation we recorded occurred inside a single session and reset at the session boundary. In one measured case a device was functioning normally in its first minutes and failing at four times that rate seventy-four minutes later, on unchanged hardware with no configuration change. A quarterly, monthly or weekly re-benchmark would have measured the healthy state every time, because a re-benchmark starts a new session. · The reassessment trigger that works is not elapsed calendar time. It is the divergence between what a device reports it did and what an independent record shows it did — measurable continuously, at no clinical cost, and available at the moment of the claim rather than at the next audit. · We have built and published that independent record. It is open source under Apache 2.0 and CDRH or any reviewer can run it. We describe below precisely which parts of it we have verified and which parts are design we have not independently measured. · For the population we build for, this is not an audit convenience. It is the restoration of a safeguard that the deployment removes. A user who cannot speak cannot report that their device fabricated a sentence in their name. Something else has to notice, and it has to notify a person who can act. FDA-2026-N-7874 — Question 19 1 The Christman AI Project 1. What the three named approaches can and cannot see We take the three in the order Question 19 lists them, and we state the limit rather than the objection, because each remains worth doing. 1.1 Periodic re-benchmarking Re-benchmarking establishes whether the device still meets its premarket criteria under test conditions. Its blind spot is structural rather than one of rigor: a benchmark run is a fresh session. Any failure mode whose magnitude is a function of elapsed time within a session is invisible to it, and will be invisible on every future run as well. Our measurement on 2026-09-03 is the concrete case. Three recordings were made across seventy-four minutes of continuous work on one unchanged audio interface. Input carrying no live signal accounted for 7.1 percent of the first file, 19.1 percent of the second, and 42.2 percent of the third. The proportion of lost input roughly doubled between each recording. Nothing was reconfigured between them. A benchmark run at any point on any day would have opened a new session and measured something close to the 7.1 percent state. 1.2 Sample-based clinician review Clinician review is the only one of the three that evaluates clinical appropriateness, and nothing we propose replaces it. Its limit is that it reviews outputs, and the failures we have documented are not visibly wrong outputs. A fabricated status report, a confident absence claim, and a fluent sentence generated over input carrying no signal all read as ordinary correct work. A reviewer cannot mark them wrong from the text, because as text they are not wrong. There is a second limit specific to sampling. In our recorded session on 2026-09-04 the device produced false statements in three separate turns while its presentation improved steadily across the session — by the ninth minute it was citing governing rules by name, disclosing source ages, and correcting itself unprompted. A sample drawn late in that trajectory scores the device as more disciplined than a sample drawn early, while the error rate is unchanged. Sampling position determines the result. 1.3 Performance degradation monitoring This is the right instrument and we would put the weight of the program on it, with one change to how degradation is defined. Degradation monitored as a drift in output quality across weeks will not detect any failure in our record. Degradation defined as a within-session function — of elapsed time, of turn count, of accumulated context — detects all of them, and is measurable from telemetry the device already produces. 2. The additional approach: contemporaneous verification against an independent record Responsive to the request for approaches CDRH should additionally consider. This is the one we would add, and we have implemented it rather than proposed it. 2.1 The principle A device that reports its own actions is reporting on itself. Where that report is the only record, it is unfalsifiable at the point of use, and eve
Original source ↗