FDA GenAI discussion / Question 12 of 26

How do you get statistically meaningful performance numbers when synthetic inputs are mixed with real ones?

Full FDA question

How should sponsors achieve statistically meaningful performance measurement for GenAI-enabled devices? Where synthetically generated inputs supplement real patient data, how should sponsors account for differences between the synthetic and real-world distributions when estimating performance, and under what conditions, if any, is it appropriate to combine benchmarking evidence and clinical confirmation evidence to support a single performance estimate?
Read the FDA discussion paper ↗

21 of 95 submissions reference this question.

All audiences
13 Industry4 Clinicians1 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/12
Filter by audience
Question 12 · Public feedback

What respondents recommend

7 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Navid Farr

Industry · Sep 8, 2026

Prespecify how performance and uncertainty are measured · Combine evidence only when justified

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
R6. Hold benchmarking to the standards CDRH already applies to bench testing (Questions 9 through 13 and 16) This follows from S4 and S6. The paper says that "test methods and acceptance criteria would be prespecified prior to testing," consistent with expectations for non-clinical bench performance testing. I strongly support this and would extend the parallel. Four points: Non-determinism. A GenAI device produces a distribution of outputs, not an output. Performance should be reported as a distribution: repeated runs on identical inputs, disclosed decoding parameters (temperature, sampling settings, seed handling), and reported variance. The paper's position in R.1 that variation in safety-critical behavior — escalation, refusal, diagnostic conclusion — is a failure rather than acceptable noise should be adopted as a firm acceptance criterion. Synthetic data lineage. Synthetic inputs generated by a model of the same class as the device under evaluation share its blind spots. They may supplement real data for stress-testing and for rare presentations, but they should never be the sole basis for a claim about subgroup performance, and the generator should be of demonstrably different lineage from the device. Sponsors should be expected to report the synthetic share of each evaluation set and to show that performance on synthetic and real inputs is concordant before combining them into a single estimate (Question 12). LLM adjudicators. Where an LLM serves as adjudicator, it should be from a different model family than the device, validated against human adjudicators on a held-out sample with reported agreement, and disclosed in the submission. Correlated error between device and judge is the obvious failure mode and it is invisible without this check. Docket No. FDA-2026-N-7874 — Individual comment — Page 6 Sequestered assets. The proposal for independent third parties to maintain sequestered evaluation datasets (Question 16) is the strongest available answer to contamination and optimization-to-the-test, and I support it. The essential safeguard is that the sponsor never sees the held-out set and cannot iterate against it. Sponsor-developed benchmarks are appropriate for demonstrating coverage of the intended use; they are not appropriate as the sole gate for authorization. On competition concerns, the ASCA model — multiple accredited bodies, published methods, FDA-recognized standards — is a reasonable template that avoids a single gatekeeper.
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

Prespecify how performance and uncertainty are measured · Report synthetic and real results separately

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 12: Statistical measurement The manufacturer must define the unit of analysis. The unit can be an output, a conversation, an encounter, a patient, an action, or a clinical result. The analysis must account for repeated patients, users, and sites. GenAI output can vary for the same input. The manufacturer must repeat tests and measure this variation. The manufacturer should report: Confidence intervals. Results for important subgroups. Results by site. Results by user type. Results by model and configuration version. Results for common and severe cases. The manufacturer should report synthetic and real data separately. Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices FDA Docket No. FDA-2026-N-7874 The manufacturer should combine the results only when it proves that the two data sources measure the same performance distribution. Benchmark results and clinical results should usually remain separate.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Combine evidence only when justified · Report synthetic and real results separately

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 12 - Statistically meaningful measurement and use of synthetic inputs FDA may wish to consider separate reporting of synthetic and real-world performance before allowing a combined estimate. A combined performance estimate is most defensible when the sponsor can justify transportability between the synthetic and real distributions, characterize weighting, and show that the combination does not obscure clinically important differences. For synthetic inputs, the evaluation record should identify the generation method, model or simulator version, source population or assumptions used to construct the synthetic cases, intended testing purpose, and the distributional relationship to the real-world target population. Real-world confirmation should remain important where clinical heterogeneity, workflow effects, or latent correlations cannot be credibly reproduced synthetically.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Prespecify how performance and uncertainty are measured

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 12 — How should statistically meaningful performance be measured? Statistically significant and clinically meaningful are not the same requirement. Demand both. Sponsors should prespecify effect sizes that matter clinically — not only ones that clear a p-value — while allowing genuinely important improvements in rare-event settings to count even when conventional statistical thresholds are difficult to reach.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Measure clinically meaningful outcomes

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 12 — Statistical methods for evaluating performance using real and synthetic data I defer to statisticians, clinical-trial methodologists, and other experts regarding the appropriate statistical design, sample sizes, confidence intervals, and methods for combining synthetic and real- world data. I would only reiterate that statistical rigor cannot compensate for choosing the wrong outcome measure. If the system is judged primarily on task completion or answer accuracy, a statistically impeccable study may still fail to measure whether the AI safely recognized urgency, latent clinical need, or appropriate escalation. Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 23 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Prespecify how performance and uncertainty are measured

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 12 — Statistically meaningful performance measurement There is a measurement problem here that classical device statistics do not address and that the paper does not raise: generative devices are non-deterministic at fixed input. The same input can produce different outputs across runs, and in some cases outputs that differ in clinical direction. Conventional performance estimation assumes a fixed input-output mapping and therefore attributes all observed variance to case mix. I recommend CDRH require: • Repeated-sampling variance reporting. Performance must be characterized across repeated runs at fixed input under the deployed sampling configuration, with within-input variance reported separately from between-input variance. • Acceptance criteria stated on a lower confidence bound. For a non-deterministic device, a point estimate meeting a threshold does not establish that the device meets the threshold. The bound, not the estimate, should be the criterion. • Sampling configuration treated as a specified device parameter subject to change control. Temperature, top-p, and equivalent settings materially change the performance distribution, and a device evaluated at one configuration and deployed at another has not been evaluated. On combining benchmarking and confirmation evidence into a single estimate: I recommend against permitting it. The two arise from different sampling frames and support different inferences — benchmarking establishes capability under constructed conditions, confirmation establishes performance under real ones. Pooling them produces an estimate that describes no actual population, and it allows large volumes of cheap benchmark data to dominate small volumes of expensive confirmation data, diluting precisely the evidence that matters most. They should be reported separately, with the confirmation estimate governing the acceptance decision and the benchmarking estimate serving as supporting evidence and as the locked postmarket baseline.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Prespecify how performance and uncertainty are measured · Combine evidence only when justified · Report synthetic and real results separately

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 12 - Statistically meaningful performance; synthetic/real evidence Trace ID. TR-Q12 | FDA Q12; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. Prespecify estimands, denominators, subgroup/trajectory strata, uncertainty intervals, and missing-data rules. Keep synthetic and real estimates separate unless distributional transportability and lineage independence justify combination. BCR rule basis. BCR-R04,R08,R11,R13,R18 Solution-stack link. S5,S7 Closure evidence. Prespecified estimands/denominators/strata/CIs; synthetic and real reported separately unless justified compatible Pass / re-open. Statistical claim maps to intended-use distribution; pooling justified when used Re-open when: Distribution shift, data- source or generator-lineage change.
Original source ↗
Source directory

All 21 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Joel GrunhutPublic / patients · Sep 7, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026