FDA GenAI discussion / Question 13 of 26

Where is synthetic data good enough, and where is it not?

Full FDA question

For which clinical domains, device functions, or subpopulations is synthetic data particularly well-suited, or particularly inadequate, as a supplement to real-world evidence? What safeguards would mitigate the risk that synthetic data generated by models of the same class as the device under evaluation reproduces the very performance gaps the evaluation is intended to detect, particularly for underrepresented subgroups?
Read the FDA discussion paper ↗

22 of 95 submissions reference this question.

All audiences
14 Industry4 Clinicians1 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/13
Filter by audience
Question 13 · Public feedback

What respondents recommend

7 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Newton’s Tree

Industry · Sep 3, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Question 13: Synthetic data Synthetic data are useful for: Scope tests. Rare but well-defined hazards. Numerical errors. Unit errors. Adversarial tests. Tool failures. Controlled changes to one patient feature. Synthetic data are not a sufficient replacement for real data about: Clinical communication. Human behavior. Long-term reliance. Local workflow. Real acquisition errors. Poorly represented groups. Clinical outcomes. A model from the same model family can reproduce the same blind spots. The manufacturer should use an independent generator where possible. The manufacturer must confirm important synthetic test results with real data.
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Response to Discussion Questions 11–13 – Clinical confirmation, representativeness, and synthetic data The level of clinical confirmation should be determined not only by criticality but also by the remaining evidentiary uncertainty after benchmarking and simulation. Prospective evidence becomes particularly important where safety depends materially on human-AI interaction, local workflow, real-time information availability, or where use of the device changes subsequent diagnostic or therapeutic decisions. In contrast, rare safety-critical failure modes may sometimes be characterized more efficiently through enriched challenge testing and simulation than through very large prospective studies. For retrospective evaluation in particular, temporal fidelity should be included as another important topic. The information provided to the device during testing should reconstruct the information state that would actually have been available at the intended clinical decision point. Later or additional investigations, diagnoses, treatment decisions, or outcome information may appropriately contribute to establishing the reference standard, but should not inadvertently be made available to the device being evaluated. Without this separation, retrospective testing may substantially overestimate performance. A related issue can arise during model development when training data contain proxy variables or information that would not be available at the intended decision point. In such cases, models may exploit shortcuts that link this additional information to the outcome. This can compromise the intended temporal and causal structure of the clinical decision problem, since model inputs should reflect information that would also be available for new cases. Representativeness should also be treated as a multidimensional concept. A dataset that accurately reflects real- world prevalence may contain too few rare but high-consequence cases to establish safety with adequate precision. Conversely, a risk-enriched challenge set is intentionally not prevalence-representative. I therefore suggest distinguishing the evidence needed to estimate expected real-world performance from the evidence needed to establish coverage of clinically and risk-relevant situations. Ultimately, the evaluation design should reflect the risks of the GenAI-enabled device as it is applied in the actual clinical setting. Statistical evaluation should correspondingly include overall performance together with prespecified subgroup, scenario, and failure-mode-specific analyses where these are relevant to safety. Sample-size requirements should be driven by the precision required for the relevant risk-based performance claims rather than by a uniform minimum number of test cases. Again, the required precision should be related to the risks in the actual clinical setting. Synthetic data can be particularly valuable for rare conditions, counterfactual testing (what-if-scenarios), controlled variation of patient or environmental characteristics, and generation of difficult conversational trajectories. They should generally augment rather than replace real data, particularly where clinical data are limited. Additionally, their regulatory role should be explicit. Synthetic data used for stress testing or coverage expansion need not reproduce actual prevalences or risk profiles exactly. Instead, synthetic data used to estimate real-world performance require substantially stronger evidence of such types of representativeness. Results from these different purposes should generally not be pooled into a single overall performance estimate. Independence of the generator, clinical plausibility, and the risk that synthetic data reproduce the same biases as the device under evaluation should also be considered. The Discussion Paper already identified the latter concern in Question 13. For devices whose safety materially depends on deployment-specific conditions, e.g., user expertise, local clinical pathways, availability of downstream safeguards, or technical integration, a proportionate form of local deployment qualification could be appropriate. This need not imply full revalidation at every institution. But the greater the regulatory reliance on local conditions, the stronger the evidence should be that those conditions actually exist and function as assumed. Response to Discussion Questions 14–15 – Reference standards and clinically relevant comparators Regarding performance comparators, I suggest more clearly separating the role of the reference standard from the role of the clinical comparator. Where a clinical comparator is used, it should reflect a relevant alternative to the device rather than being conflated with the reference standard. According to standard rules for medical devices, the GenAI device needs to be compared to the performance of the established standard-of-care to gain market access. This means, that it needs to be assessed how the GenAI system performs in relation to the standard-of-care when comparing both outcomes to an idealized reference standard or ground truth. The reference standard itself should establish, as reliably and accurately as possible, what diagnosis, assessment, or course of action is correct or clinically appropriate for the particular test case. Instead, the comparator, i.e. standard-of-care, should represent the care that would realistically occur in the absence of the device, including the intended user group and clinical environment. The Discussion Paper already raises this possibility in Question 15, including unaided clinical judgment, delayed specialist review, or no intervention. Accordingly, the experts establishing the reference standard need not be the clinicians against whom the device is compared. For a device intended to support generalist physicians, for example, specialist adjudication may establish the reference standard while representative generalist physicians provide the clinically releva
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Question 13 - Appropriate and inappropriate uses of synthetic data Synthetic data may be particularly useful for rare-event stress testing, controlled perturbations, privacy-preserving development of standardized scenarios, device/tool failure simulations, and targeted testing of under-sampled edge cases. It is less reliable when the clinically important phenomenon depends on complex latent relationships, undocumented workflow behavior, or population characteristics that the generator does not represent well. A specific safeguard is needed when synthetic data are generated by models of the same or closely related class as the device under evaluation. Correlated blind spots can create false reassurance. Mitigations include independent generation methods, real-world holdout confirmation, explicit provenance, subgroup-specific distribution checks, adversarial challenge cases, and independent clinical adjudication. Independent clinical adjudication is particularly useful because it introduces an evaluation reference that does not necessarily share the same model-class assumptions or learned failure modes as either the device or the synthetic-data generator. Synthetic evidence should generally be treated as complementary to, rather than a replacement for, evidence of real-world diversity unless transportability is well established.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Question 13 — Where should synthetic data be used? Synthetic data should expand the test universe. It should never become the evidence universe. Synthetic data is valuable for rare events, adversarial scenarios, and situations too dangerous to induce in real patients. But synthetic data generated by models similar to the one under evaluation can reproduce the same blind spots it is meant to catch. FDA should favor independently generated and validated synthetic datasets for any decision with meaningful clinical consequence.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Question 13 — Where is synthetic data useful, and where might it be inadequate? Response Synthetic data can be extremely useful for expanding testing, constructing rare scenarios, protecting privacy, and deliberately challenging systems with high-risk situations that would be difficult or unethical to recreate prospectively. But synthetic patients can also become too rational, too complete, and too clinically tidy. That is especially dangerous in patient-facing AI evaluation. A synthetic case generator may create the classic textbook presentation of myocardial infarction. The real patient may say: “I ate too much pizza and have awful heartburn. Can you send me my insurance card?” If synthetic testing reproduces what clinicians expect patients to say rather than how patients actually behave, the benchmark may systematically miss the very failures most likely to harm people. Synthetic evaluation should therefore deliberately model:  incomplete disclosure;  symptom minimization;  contradictory information;  low health literacy;  financial concern;  avoidance;  emotional distress;  culturally varied descriptions of symptoms; Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 21 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  multimorbidity;  atypical presentation;  and administrative proxy questions concealing clinical need. FDA should also consider whether synthetic datasets reproduce the biases and assumptions of the models that generate them. Synthetic data can strengthen safety evaluation. It should not replace genuine real-world human behavior.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
Question 13 — Synthetic data The paper’s framing of this question — the risk that synthetic data generated by models of the same class as the device reproduces the very gaps evaluation should detect — is exactly right, and I want to press on the implication. The concern is not bias in the usual sense; it is correlated failure. A generator sharing training data, architecture, or lineage with the device under test shares its blind spots. Where the device misunderstands a rare presentation, a same-family generator will produce synthetic cases embodying the same misunderstanding. The evaluation then returns a clean result on a population that does not exist, and the failure is invisible because the evidence looks complete. This is worse than having no synthetic data, because it manufactures unwarranted confidence. 8 of 19 Docket No. FDA-2026-N-7874 I recommend CDRH draw a bright line by purpose rather than by domain: • Synthetic data is appropriate for finding failures. Stress testing, adversarial probing, coverage of rare and dangerous scenarios that cannot be ethically or practically collected. Here a failure discovered is informative regardless of whether the input was real, and the correlated-blind-spot problem biases toward under-detection — a conservative direction. • Synthetic data is not appropriate for establishing performance rates. Any rate requires a denominator representing a real population. Synthetic data cannot supply one, because the generator’s distribution is unvalidated against the deployment distribution and unvalidatable except by reference to the real data one is trying to avoid collecting. I further recommend that same-family generation be prohibited for subgroup performance claims specifically. Subgroup performance is where correlated blind spots do the most damage — underrepresented subgroups are underrepresented in the generator’s training data for the same reasons they are underrepresented in the device’s — and it is where a false negative has the clearest equity consequence. Where synthetic data is used at all, the generator’s identity, provenance, and training lineage should be disclosed, and its relationship to the device’s underlying model stated explicitly. I am not aware of a validated method for establishing that a synthetic clinical population is distributionally adequate to a real one for regulatory purposes, and would be interested to see any that commenters can point to. I would rather CDRH proceed as though none exists than adopt a permissive posture on the assumption that one will emerge.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Use synthetic cases for rare events and stress testing · Check for shared blind spots in generated test data · Keep real evidence for claims synthetic data cannot establish

This filing recommends: use synthetic cases for rare events and stress testing; check for shared blind spots in generated test data; keep real evidence for claims synthetic data cannot establish. The passage gives the applicable scope and conditions.

Read the source passage
FDA Question 13 - Where synthetic data helps or fails Trace ID. TR-Q13 | FDA Q13; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. Use synthetic data for combinatorial stress, rare known scenarios, perturbations, and privacy-constrained coverage; do not rely on it alone for unknown blind spots or safety-critical claims where generator and device may share correlated gaps. Add independent real sentinels and lineage disclosure. BCR rule basis. BCR-R05,R11,R13,R15 Solution-stack link. S5,S7 Closure evidence. Generator lineage, independent real sentinels, diverse generators, subgroup stress tests Pass / re-open. Synthetic evidence supplements independent real witnesses where blind spots are plausible Re-open when: Generator/model lineage change or new subgroup gap.
Original source ↗
Source directory

All 22 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Joel GrunhutPublic / patients · Sep 7, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026