FDA GenAI discussion / Question 10 of 26

How can a benchmark score be shown to predict real-world behavior?

Full FDA question

CDRH recognizes that publicly available benchmarking assets may be subject to data contamination, saturation, and limited real-world representativeness. How should a sponsor establish that performance on a given benchmark predicts safe and effective real-world behavior for the device’s intended use? What evidence should support the construct validity of a benchmark used to gate device evaluation, and what role should sponsor-developed benchmarks play given potential concerns around independence and optimization to the test?
Read the FDA discussion paper ↗

26 of 95 submissions reference this question.

All audiences
16 Industry5 Clinicians2 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/10
Filter by audience
Question 10 · Public feedback

What respondents recommend

11 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Sam Rosenthal (Red Kit)

Industry · Sep 9, 2026

Test realistic and challenging clinical scenarios

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 — benchmark validity and sponsor-developed benchmarks. Three things I have learned building my own scenario suite, offered for what they are worth: • Mechanical scanning is not a release check. An automated pass that checks an output for required and forbidden content (the right steps present, the wrong technique absent) can pass a set of outputs that a careful human read then fails, because the human notices that the answer is correct for a different premise than the one the user described. Any benchmark used to gate a lay-rescuer function should include premise-aware items — scenarios whose opening line changes which protocol is correct — and should require human adjudication on those items, not keyword or rubric scoring alone. • Age-stratified negatives. The suite must include items where the adult-correct technique is the infant-wrong one, scored so that the adult answer fails. A benchmark built from adult scenarios will not detect a model that has learned one template. • Sponsor-developed benchmarks are unavoidable and should be disclosable. There is no public benchmark for bystander first aid on a phone. I would support a requirement that a sponsor- developed benchmark be disclosed in full — items, rubric, adjudication protocol, and the model's outputs — so that independence can be checked by anyone, rather than a requirement that the benchmark itself be independent, which would leave small developers with no benchmark at all.
Original source ↗

Bhasker Sambar, M.Pharm.

Industry · Sep 4, 2026

Protect test sets from exposure or contamination

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 — Construct validity, contamination, and saturation Contamination is the biggest concern. If a model was trained on public benchmark questions, it may score well because it has seen the answers before, not because it can reason through the task. Sponsors may not be able to tell, because they usually cannot see the full training data. Three controls can help now: – Provenance dating — compare each benchmark item’s publication date with the model’s training cutoff. Items that came before the cutoff should be treated as potentially contaminated unless the sponsor can show otherwise. – Canary items — include held-out questions designed so memorization gives a different answer than genuine reasoning, such as questions with changed numerical values. – Sequestered holdout — for higher-risk functions, include some testing on data the sponsor has not seen before. Sponsor-developed benchmarks should be allowed, and in many cases they will be necessary because public benchmarks may not match a narrow intended use. The safeguard should be a clear, prespecified benchmark plan, not a ban. Sponsors should explain how the benchmark was built before generating results, similar to a prospectively defined non-clinical protocol. Recommendation. Require provenance dating and canary items when public benchmarks are used to support a regulatory decision. For sponsor-developed benchmarks, require a prespecified construction plan. For higher-risk functions, require some evaluation using sequestered holdout data.
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Use independent testing or test custodians · Protect test sets from exposure or contamination

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10: Benchmark construct validity. Sponsor-developed benchmarks are valuable for intended-use specificity and should be permitted, but they should be complemented by at least one sequestered test set held by a party structurally independent of both the sponsor and the foundation model developer, to guard against optimization to the test, a role naturally suited to the third-party mechanisms discussed in Question 16.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Use independent testing or test custodians · Protect test sets from exposure or contamination · Check results against real-world clinical evidence

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 - Benchmark contamination, saturation, and construct validity A benchmark should be treated as evidence only to the extent that its construct validity for the intended use is established. Sponsors should justify why performance on the benchmark is expected to predict clinically relevant behavior in the deployment environment. Useful safeguards include prespecified evaluation plans, sequestered or independently maintained test sets where practical, disclosure of known contamination risks [12], testing on out-of-distribution and adversarial cases, evaluation across clinically relevant settings and subgroups, and confirmation on real or clinically representative inputs. Sponsor- developed benchmarks can be valuable when the intended use is specialized, but their design and acceptance criteria should be transparent enough to reduce optimization-to-the- test risk. 7
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Use independent testing or test custodians · Protect test sets from exposure or contamination · Check results against real-world clinical evidence

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 — Do benchmarks predict real-world performance? A benchmark a developer can optimize against isn't evidence. It's a target. Public benchmarks support transparency but are insufficient alone, since developers can train toward them. I recommend a three-layer evidentiary structure: public benchmarks for comparison, sequestered regulatory benchmarks for independent assessment, and real-world clinical confirmation to demonstrate that benchmark performance actually translates to the clinic. No single layer should be sufficient for high-risk AI.
Original source ↗

VivaSecuris

Industry · Aug 25, 2026

Protect test sets from exposure or contamination · Check results against real-world clinical evidence

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10. Benchmark validity should be demonstrated through an evidence package, not inferred from popularity or score stability. A sponsor should document the benchmark’s intended construct, clinical relevance, coverage, known gaps, contamination analysis, subgroup representation, scoring reliability, adjudicator agreement, and relationship to real-world outcomes. Sponsor-developed benchmarks can be valuable for narrow intended uses, but high-consequence claims should include sequestered or independently governed assets to reduce test optimization and leakage.
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Protect test sets from exposure or contamination

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 raises data contamination, saturation, and limited real-world representativeness in publicly available benchmarking assets. The concern is acute for the Safety elements specifically. Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 6 Because the provenance of training, instruction-tuning, and safety-tuning corpora is generally incomplete or undisclosed, public adversarial benchmarks should be treated as potentially exposed rather than presumed sequestered. A model may decline a prompt drawn from a well-known adversarial dataset because that prompt or a near neighbor was present during alignment training, rather than because the underlying behavior generalizes. Measured resistance on public adversarial assets therefore tends to overstate resistance to novel constructions within the same attack class, and the overstatement grows as an asset ages and circulates. This interacts with Question 16. Sponsor-developed adversarial assets are less likely to be exposed but carry an evident independence problem, since the same party selects both the attacks and the acceptance criteria. One available balance is to publish the attack category taxonomy and scoring methodology while retaining the specific probe payloads, which preserves reviewability of the construct without publishing a reproduction recipe. Sequestered assets held by a qualified independent party would address both concerns more completely, at the cost of the program design considerations raised in Question 16. Recommendation: For the S-series elements, a sponsor's construct validity argument should address exposure specifically, including whether assets post-date the model's training data and whether they appear in public alignment datasets. Where sponsor-developed adversarial assets are used, the taxonomy of attack classes and the scoring rubric should be prespecified and disclosed even where individual payloads are not. 8. Agentic systems: tie acceptance criteria to the action surface (Question 26) Element A.1 appropriately includes resistance to prompt injection through user inputs, retrieved content, and tool outputs, which reflects the established finding that indirect injection through retrieved content is a distinct attack surface from direct user input (Greshake et al., AISec 2023). I would add one consideration about how the elements interact for agentic devices. A boundary failure under S.2 in a non-agentic informational device produces an inappropriate output that a user may or may not rely upon. The same failure in a device with tool access produces an action. The consequences axis in Figure 1 captures the severity of relying on an incorrect output, but for agentic systems the relevant quantity is closer to the severity of an action taken without any opportunity for reliance to be withheld. Recommendation: For devices in the action-taking columns, acceptance criteria should account not only for the likelihood of unsafe generation but also for the action surface exposed to the model, the reversibility of available actions, the authorization scope granted to the device, the availability of independent confirmation before high-consequence or irreversible actions, and the capability to halt or roll back an in-progress action sequence. Where meaningful autonomous action is possible, S.2 and A.1 should be evaluated as jointly interacting controls rather than as independent checklist items, since the human review that moderates risk elsewhere in the framework is absent by construction. Closing The discussion paper is correct that the range of possible inputs to a GenAI-enabled device may be too large for exhaustive testing to be practical, and the competency-based structure is a sound response to that constraint. My comments concern the durability of that structure across the device lifecycle: that the artifact evaluated at Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 7 premarket is the artifact that reaches the patient, that changes to it are enumerated in a way that prompts characterization, that the safety elements are re-measured actively rather than inferred from observational use, and that results are compared at a resolution fine enough to make a regression visible. I would be glad to provide further detail on any of the above if it would be useful to the Center. Respectfully submitted, Polina Moshenets Founder, SichGate Polina.Moshenets@sichgate.com Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 8
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Test realistic and challenging clinical scenarios

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 — How can benchmarking predict real-world behavior? Response Benchmarks need messy humans. Real patients do not interact with healthcare technology like board-examination questions. They use slang. They omit facts. They misunderstand terminology. They minimize symptoms. They become embarrassed. They change their minds. They worry about money. They forget information. They rationalize delay. They ask proxy questions. They sometimes resist exactly the advice most likely to protect them. A benchmark populated primarily with clinically explicit, complete, rational prompts will substantially overestimate the safety of patient-facing AI. PatientAgentBench is valuable precisely because it begins to expose that distinction. Its sustained conversations and tool-use scenarios reveal failures that static question-and-answer benchmarks can miss. FDA should encourage benchmark designs that intentionally evaluate:  indirect presentation;  incomplete information;  symptom minimization;  clinically complex patients asking routine administrative questions;  multimorbidity;  financial barriers;  health literacy;  emotional distress;  cognitive bias;  patient reluctance;  escalation resistance;  tool use;  longitudinal context; and  the difference between successfully executing an action and safely managing an encounter. Benchmarks should also test for a particularly dangerous form of AI failure: premature closure after apparent success. In many industries, successful completion of the requested task marks the end of the interaction. In healthcare, that may be exactly the wrong moment to stop. The question is not merely: “Did the AI complete the task?” The question is: “Did it know whether completing the task was enough?” IV. Appropriate Clinical Comparator and Evidence Standard
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Protect test sets from exposure or contamination

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 10 — Benchmark validity, contamination, and sponsor-developed benchmarks This question deserves the most attention of any in the paper, because benchmarking is load-bearing for the entire premarket structure and the assets currently available cannot bear that load. Contamination is not testable without a training-data cutoff attestation. A sponsor cannot demonstrate that a public benchmark was absent from a third-party model’s training corpus without information the sponsor does not possess. I recommend CDRH state that where a device relies on a third-party model, benchmarking evidence based on publicly available assets is not interpretable absent a dated training-data cutoff attestation from the model provider, and that temporally held-out evaluation data — constructed from material post-dating the attested cutoff — is the preferred mitigation. This also creates the market pull discussed under Question 25. 6 of 19 Docket No. FDA-2026-N-7874 Construct validity must be argued, not assumed. I recommend CDRH require a sponsor to state, for each benchmark used, the specific real-world behavior it is intended to predict, the reasoning connecting them, and the known limitations of that connection. This is a familiar requirement in a different vocabulary: it is analytical validation, and it should be documented to the same standard. Sponsor-developed benchmarks should be permitted but structurally separated. They are often unavoidable — for a novel intended use, no external benchmark exists. The controlling concern is not that the sponsor built the benchmark but that the sponsor may have optimized against it. I recommend three safeguards: (a) the benchmark and its acceptance thresholds are locked and submitted before evaluation, not after; (b) the sponsor attests, subject to audit, that the gating set was not used for model selection, prompt engineering, fine-tuning, or any iteration on the device — the separation between development and gating sets must be demonstrable from development records, not merely asserted; (c) a held-out portion is escrowed with FDA or a recognized third party and never returned to the sponsor. Public benchmarks should be characterized as screening, not evidence. Given contamination and saturation, strong performance on a public benchmark establishes little; weak performance establishes a great deal. I recommend CDRH treat public benchmarks as necessary-but-not-sufficient screens and state explicitly that they do not constitute evidence of effectiveness. A rotating non-public evaluation resource is worth serious consideration. The contamination problem is structurally unsolvable by sponsors acting individually, since any benchmark that becomes valuable becomes public and then becomes training data. A periodically refreshed evaluation set, maintained by FDA or a recognized body and never published, is one of the few mechanisms that addresses the root cause. I recognize this carries substantial resource and governance implications, and I raise it as a direction worth evaluating rather than as a costed proposal.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Use independent testing or test custodians · Protect test sets from exposure or contamination · Check results against real-world clinical evidence

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 10 - Benchmark construct validity Trace ID. TR-Q10 | FDA Q10; Sec. V.E; App. B; pp. 18-19 / 28-29 BCR response. Require evidence that benchmark performance predicts intended-use behavior: independent/sequestered assets, contamination controls, representative morphology, subgroup/trajectory coverage, perturbation stability, and correlation with clinical/postmarket witnesses. BCR rule basis. BCR-R05,R08,R11,R15,R16 Solution-stack link. S5,S7,S12 Closure evidence. Sequestered assets, contamination checks, perturbation stability, independent adjudication, clinical/postmarket linkage Pass / re-open. Benchmark predicts intended-use behavior within defined envelope Re-open when: Benchmark exposure/saturation, deployment distribution, model revision.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Protect test sets from exposure or contamination

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 10: Public Benchmark Contamination — A Serious Problem Requiring a Structural Solution Public benchmark contamination is not a theoretical concern. It is a documented phenomenon in the AI industry: widely used public benchmarks have been incorporated — intentionally or inadvertently — into training datasets, rendering them unreliable as independent performance measures. For medical device applications, this is not merely an academic problem. A device that appears to perform well on a contaminated benchmark may fail in clinical practice in ways that could harm patients. FDA should establish a sequestered national benchmark repository for medical AI, developed in collaboration with NIST and maintained by an independent steward, analogous to NIST's existing machine learning benchmarking challenges. Access to the sequestered evaluation set should be tightly controlled, with submission and scoring handled through a blinded third-party process. Sponsor-developed benchmarks should require third-party validation of construct validity before FDA will accept them as primary evidence. The current reliance on publicly available benchmarks, without structural controls against contamination, is a vulnerability that adversarial actors could exploit and that well-intentioned manufacturers may inadvertently fall into.
Original source ↗
Source directory

All 26 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026SichGate Inc.Industry · Aug 22, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026