Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish
This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.
Read the source passage
Response to Discussion Questions 11–13 – Clinical confirmation, representativeness, and synthetic data The level of clinical confirmation should be determined not only by criticality but also by the remaining evidentiary uncertainty after benchmarking and simulation. Prospective evidence becomes particularly important where safety depends materially on human-AI interaction, local workflow, real-time information availability, or where use of the device changes subsequent diagnostic or therapeutic decisions. In contrast, rare safety-critical failure modes may sometimes be characterized more efficiently through enriched challenge testing and simulation than through very large prospective studies. For retrospective evaluation in particular, temporal fidelity should be included as another important topic. The information provided to the device during testing should reconstruct the information state that would actually have been available at the intended clinical decision point. Later or additional investigations, diagnoses, treatment decisions, or outcome information may appropriately contribute to establishing the reference standard, but should not inadvertently be made available to the device being evaluated. Without this separation, retrospective testing may substantially overestimate performance. A related issue can arise during model development when training data contain proxy variables or information that would not be available at the intended decision point. In such cases, models may exploit shortcuts that link this additional information to the outcome. This can compromise the intended temporal and causal structure of the clinical decision problem, since model inputs should reflect information that would also be available for new cases. Representativeness should also be treated as a multidimensional concept. A dataset that accurately reflects real- world prevalence may contain too few rare but high-consequence cases to establish safety with adequate precision. Conversely, a risk-enriched challenge set is intentionally not prevalence-representative. I therefore suggest distinguishing the evidence needed to estimate expected real-world performance from the evidence needed to establish coverage of clinically and risk-relevant situations. Ultimately, the evaluation design should reflect the risks of the GenAI-enabled device as it is applied in the actual clinical setting. Statistical evaluation should correspondingly include overall performance together with prespecified subgroup, scenario, and failure-mode-specific analyses where these are relevant to safety. Sample-size requirements should be driven by the precision required for the relevant risk-based performance claims rather than by a uniform minimum number of test cases. Again, the required precision should be related to the risks in the actual clinical setting. Synthetic data can be particularly valuable for rare conditions, counterfactual testing (what-if-scenarios), controlled variation of patient or environmental characteristics, and generation of difficult conversational trajectories. They should generally augment rather than replace real data, particularly where clinical data are limited. Additionally, their regulatory role should be explicit. Synthetic data used for stress testing or coverage expansion need not reproduce actual prevalences or risk profiles exactly. Instead, synthetic data used to estimate real-world performance require substantially stronger evidence of such types of representativeness. Results from these different purposes should generally not be pooled into a single overall performance estimate. Independence of the generator, clinical plausibility, and the risk that synthetic data reproduce the same biases as the device under evaluation should also be considered. The Discussion Paper already identified the latter concern in Question 13. For devices whose safety materially depends on deployment-specific conditions, e.g., user expertise, local clinical pathways, availability of downstream safeguards, or technical integration, a proportionate form of local deployment qualification could be appropriate. This need not imply full revalidation at every institution. But the greater the regulatory reliance on local conditions, the stronger the evidence should be that those conditions actually exist and function as assumed. Response to Discussion Questions 14–15 – Reference standards and clinically relevant comparators Regarding performance comparators, I suggest more clearly separating the role of the reference standard from the role of the clinical comparator. Where a clinical comparator is used, it should reflect a relevant alternative to the device rather than being conflated with the reference standard. According to standard rules for medical devices, the GenAI device needs to be compared to the performance of the established standard-of-care to gain market access. This means, that it needs to be assessed how the GenAI system performs in relation to the standard-of-care when comparing both outcomes to an idealized reference standard or ground truth. The reference standard itself should establish, as reliably and accurately as possible, what diagnosis, assessment, or course of action is correct or clinically appropriate for the particular test case. Instead, the comparator, i.e. standard-of-care, should represent the care that would realistically occur in the absence of the device, including the intended user group and clinical environment. The Discussion Paper already raises this possibility in Question 15, including unaided clinical judgment, delayed specialist review, or no intervention. Accordingly, the experts establishing the reference standard need not be the clinicians against whom the device is compared. For a device intended to support generalist physicians, for example, specialist adjudication may establish the reference standard while representative generalist physicians provide the clinically releva
Original source ↗