Use independent parties to hold or maintain test assets · Control conflicts and keep evaluation open to competition
Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.
Read the source passage
R6. Hold benchmarking to the standards CDRH already applies to bench testing (Questions 9 through 13 and 16) This follows from S4 and S6. The paper says that "test methods and acceptance criteria would be prespecified prior to testing," consistent with expectations for non-clinical bench performance testing. I strongly support this and would extend the parallel. Four points: Non-determinism. A GenAI device produces a distribution of outputs, not an output. Performance should be reported as a distribution: repeated runs on identical inputs, disclosed decoding parameters (temperature, sampling settings, seed handling), and reported variance. The paper's position in R.1 that variation in safety-critical behavior — escalation, refusal, diagnostic conclusion — is a failure rather than acceptable noise should be adopted as a firm acceptance criterion. Synthetic data lineage. Synthetic inputs generated by a model of the same class as the device under evaluation share its blind spots. They may supplement real data for stress-testing and for rare presentations, but they should never be the sole basis for a claim about subgroup performance, and the generator should be of demonstrably different lineage from the device. Sponsors should be expected to report the synthetic share of each evaluation set and to show that performance on synthetic and real inputs is concordant before combining them into a single estimate (Question 12). LLM adjudicators. Where an LLM serves as adjudicator, it should be from a different model family than the device, validated against human adjudicators on a held-out sample with reported agreement, and disclosed in the submission. Correlated error between device and judge is the obvious failure mode and it is invisible without this check. Docket No. FDA-2026-N-7874 — Individual comment — Page 6 Sequestered assets. The proposal for independent third parties to maintain sequestered evaluation datasets (Question 16) is the strongest available answer to contamination and optimization-to-the-test, and I support it. The essential safeguard is that the sponsor never sees the held-out set and cannot iterate against it. Sponsor-developed benchmarks are appropriate for demonstrating coverage of the intended use; they are not appropriate as the sole gate for authorization. On competition concerns, the ASCA model — multiple accredited bodies, published methods, FDA-recognized standards — is a reasonable template that avoids a single gatekeeper.Original source ↗