Use that comparator, with conditions
Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.
Read the source passage
Discussion Questions 14–15 – Reference standards and clinically relevant comparators Regarding performance comparators, I suggest more clearly separating the role of the reference standard from the role of the clinical comparator. Where a clinical comparator is used, it should reflect a relevant alternative to the device rather than being conflated with the reference standard. According to standard rules for medical devices, the GenAI device needs to be compared to the performance of the established standard-of-care to gain market access. This means, that it needs to be assessed how the GenAI system performs in relation to the standard-of-care when comparing both outcomes to an idealized reference standard or ground truth. The reference standard itself should establish, as reliably and accurately as possible, what diagnosis, assessment, or course of action is correct or clinically appropriate for the particular test case. Instead, the comparator, i.e. standard-of-care, should represent the care that would realistically occur in the absence of the device, including the intended user group and clinical environment. The Discussion Paper already raises this possibility in Question 15, including unaided clinical judgment, delayed specialist review, or no intervention. Accordingly, the experts establishing the reference standard need not be the clinicians against whom the device is compared. For a device intended to support generalist physicians, for example, specialist adjudication may establish the reference standard while representative generalist physicians provide the clinically relevant comparator. Similarly, for patient-facing home-use devices, representative patients or lay users may be the relevant user comparator, while the reference standard may still rely on specialist adjudication or another high-quality clinical reference. This approach shifts the focus from an absolute comparison of GenAI with an expert, who may not represent the intended user, toward a relative assessment in which the GenAI-enabled device and the intended user group are compared by reference to the same reference standard. Based on this, the demonstration of non-inferiority to the comparator as a relative criterion gets the main objective instead of determining an absolute level of deviation between the GenAI and the reference standard. The evaluation should also generally focus on the combined human-AI system rather than comparing the stand- alone GenAI-enabled device with the intended user. More generally, the following clinical comparison may be appropriate: • intended user + standard of care + GenAI device versus • intended user + standard of care without the GenAI device, with both evaluated against the same independent reference standard. This provides a more direct assessment of whether the device improves or at least preserves decision quality within its intended context of use than a simple comparison of stand-alone GenAI output with expert opinion. Relative non-inferiority or superiority approaches may help avoid arbitrary absolute performance thresholds, although critical safety outcomes should still be subject to absolute risk-based acceptance criteria. This approach also creates a clearer distinction between the information used to establish the reference standard and the information available to the evaluated users or device. While subsequent diagnostic findings or follow-up may strengthen the reference standard, the investigational and comparator conditions should only have access to information that would realistically have been available at the relevant decision point. Overall conclusion for Section V In summary, I support the proposed combination of competency-based benchmarking and clinical confirmation, but suggest strengthening the framework, in particular, in the following areas. • better discriminate between evaluating errors in the output of the GenAI device and consequences that result from these errors; • more explicitly link evidence intensity to criticality while using Product-Specific Risk Management to determine the specific content of evaluation; • expand benchmarking from predominantly dataset-based assessment toward dynamic, risk-based evaluation suites; • establish the use of an independent reference standard where the GenAI device can be compared to the established standard of care in a relative way; • distinguish intrinsic device performance from performance of the device within the intended human and clinical system; • treat evaluation as a lifecycle process in which critical assumptions, safeguards, and the authorized operating envelope can be progressively confirmed, expanded, and periodically reassessed. Such an approach would preserve the scalability sought by a competency-based framework while providing a clearer link between device competence, risk-control effectiveness, clinical decision quality, and ultimately reasonable assurance of safety and effectiveness. Section VI – Postmarket Mon
Original source ↗