← All 123 filings

Alfredo Di Giovanni

CliniciansClinicianFiled September 26, 20263,921 words · 1 attachmentFDA-2026-N-7874-0123
LinkedInX

Themes it raises

8 of the 21 themes in the docket, each with the passage we counted, verbatim.
Judging devices the way clinicians are credentialedFDA Q7, Q8
“I support the competency-based approach. When inputs and outputs cannot be enumerated, competence cannot be verified exhaustively; it must be inferred from a sample of performance.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Licensing boards address analogous contamination and test- preparation problems with secure item banks, regularly retired and refreshed items, and statistically linked forms.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“The case is the unit of inference. Repeated runs of one case estimate within-case variability; they add no information about performance across cases.”
Whether the user can judge the outputFDA Q3, Q4
“The choice between generalist and specialist should follow the intended use: when a generalist uses the device on a specialist question, safety should be judged against the specialist standard.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“For clinician-facing devices the main harm pathway is automation bias (Goddard et al., 2012). I recommend reporting an error-propagation rate: the proportion of incorrect device outputs that clinicians adopt.”
Who is accountable when something goes wrongFDA Q21
“What can be credentialed is the socio-technical system — device, scope, supervising clinicians, monitoring — with responsibility allocated explicitly among the manufacturer, the deploying institution and the supervising clinicians (see also Q21).”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“Evaluation should be version-pinned, and contracts with model providers should guarantee access to pinned versions and notice of deprecation long enough to complete the appropriate tier (see also Q25).”
Watching the device after it shipsFDA Q19, Q20
“Between formal re-examinations, a small set of safety-critical cases with known expected behavior can be run against the deployed configuration on a fixed schedule.”

FDA questions it names

Questions this filing names by number.

Q7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ12 · Statistically meaningful performanceQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ24 · Third-party foundation model changesQ25 · Foundation Model Master Files

Machine-assisted draft, pending human review. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Please find attached my comment on FDA Docket FDA-2026-N-7874, addressing selected questions concerning the competency-based evaluation of generative AI-enabled medical devices. The attached document contains the full comment and supporting references.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Comment on Docket FDA-2026-N-7874
September 26, 2026
To: Dockets Management Staff (HFA-305), U.S. Food and Drug Administration.
Re: Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and
Request for Feedback (August 18, 2026).
From: Alfredo Di Giovanni, MD, Ophthalmology Service, ASL Napoli 2 Nord, Naples, Italy.

The competency-based approach is the right frame for generative AI-enabled devices. It transfers the
structure of professional credentialing, but not the measurement properties that make credentialing
trustworthy. This comment proposes how to re-establish those properties for a non-human, stochastic
examinee. The rigor and depth of the recommendations below should be scaled to intended use, device
activity and the consequences of an incorrect output under the two-axis risk framework (Q8).

Summary of recommendations
• Complement software validation; do not replace it (Q7). Scope every competency claim to a task,
population, oversight level and version, as entrustment decisions are scoped in clinical training.
• Add a benchmarking element for user-conditioned drift (Q9). A new element, R.3 (invariance to
user state and expressed belief), would test whether clinical judgment is unaffected by account
history, persistent memory, personalization and the user's expressed beliefs or authority. Pair it with
appropriate corrigibility: updating on new evidence, not on assertion.
• Specify perturbations in both directions (Q9). Require a clinician-prespecified taxonomy that
separates perturbations that must not change the output from those that must.
• Validate the examination before the examinee (Q10). Require a validity argument for any gating
benchmark, organized on the five sources of validity evidence in the Standards for Educational and
Psychological Testing. It should include evidence linking benchmark scores to clinical-confirmation
outcomes.
• Secure the item bank (Q10). Prefer sequestered, independently governed item banks with periodic
refresh and statistically linked forms. Where no independent benchmark exists, allow a sponsordeveloped gating benchmark only if it is prespecified, an independent custodian holds its final items
and its scoring or adjudication is independent.
• Make the case the unit of inference (Q12). Report variance components across cases, runs,
paraphrases, contexts and adjudicators. Prespecify cases and runs per case, and apply zero-failure
bounds to safety-critical behavior at both levels.
• Treat the reference panel as an instrument (Q14). Report its reliability, define sets of acceptable
answers for open-ended outputs, and set thresholds with established standard-setting methods.
Safety-critical elements should use conjunctive standards that proficiency cannot offset.
• Evaluate the team when a clinician is in the loop (Q14). Base the primary criterion on human–AI
team performance, including how often clinicians adopt incorrect outputs.
• Test whether the image drives the answer (Q17). For vision–language devices, require imageablation and image-substitution tests. Include cases in which color-coded normative classifications
diverge from the correct interpretation.

Page 1 of 8
• Treat each foundation-model version as presumptively a new candidate (Q22, Q24). Scale reexamination to the behavioral change actually observed, from sentinel testing to full re-examination
on linked forms, not to the nominal size of the update.

Q7 — Overall position
I support the competency-based approach. When inputs and outputs cannot be enumerated,
competence cannot be verified exhaustively; it must be inferred from a sample of performance.
That is
how clinicians are credentialed, and it is the right logic for devices that produce open-ended clinical
judgments.
The case for this approach does not depend on whether a model "reasons" in a human sense. It rests on
the structure of the output: an open-ended recommendation that someone will act on, with no single
reference label to score it against. Output of that form should be evaluated as a professional judgment
is evaluated.
Complement, not replacement. Cybersecurity, version control, logging, data integrity and failure-mode
analysis remain necessary. They are not sufficient, because none of them shows that the device is
clinically competent.
Structure versus measurement properties. The analogy transfers the structure of credentialing:
examinations, supervised practice, scoped privileges, recertification. It does not transfer the
measurement properties that make credentialing trustworthy. For human candidates, decades of
psychometric work establish how many cases a valid inference needs, how reliable examiners are and
how passing standards are set. None of this can be assumed for a non-human, stochastic, contextsensitive examinee; it must be re-established.
In one respect a device is easier to examine than a person. It can be re-examined on the same case,
under controlled counterfactual variation, with earlier attempts kept out of its context. That control
exists only if evaluation fixes and discloses memory and personalization settings, which I address under
Q9.
Scope. In competency-based medical education, entrustment is granted per activity, not per person: a
resident may be entrusted with phacoemulsification long before vitreoretinal surgery (ten Cate, 2013).
The corresponding claim for a device reads: version V, in configuration C, is competent for task T, in
population P, at oversight level O. The discussion paper's activity axis maps closely onto the supervision
levels of entrustment scales. Moving a device along that axis should require new evidence, as it does for
a trainee.
Accountability. I share the concern of Freyer et al. (2025), cited in the discussion paper, that the analogy
weakens at accountability. A device can be disabled, withdrawn or recalled, but it does not itself bear
professional or reputational consequences; accountability remains with the manufacturer and, where
applicable, the deploying institution. What can be credentialed is the socio-technical system — device,
scope, supervising clinicians, monitoring — with responsibility allocated explicitly among the
manufacturer, the deploying institution and the supervising clinicians (see also Q21).

Q9 — A missing element: user-conditioned drift
Element R.1 already covers repeated runs, paraphrases, the order of information and long
conversations. It does not explicitly address state that persists across sessions, or deference to what the

Page 2 of 8
user asserts. Both bend the device's judgment toward the user rather than the patient; I call this userconditioned drift.
Cross-session adaptation. In earlier work I described how repeated use within one diagnostic field may
induce uncontrolled semantic adaptation: the model aligns with the patterns that dominate the user's
earlier interactions (Di Giovanni, 2025). The same query to the same model may then perform
differently depending on the account's history. Controlling history, as that work recommended for
research studies, isolates capability. Regulatory evaluation must also test the conditioned state that
routine use produces through memory and personalization features. The risk is greatest where inputs
are ambiguous, such as borderline imaging findings, because context has the most room to decide the
reading.
Deference to the user. Models can follow the user's framing against their own knowledge. Sharma et al.
(2024) showed across five AI assistants that models can shift toward users' stated views and wrongly
admit mistakes when challenged. In a medical setting, GPT-family models complied with 100% (50/50) of
illogical requests to produce false drug-equivalence information (Chen et al., 2025). The proposed
elements cover users who minimize symptoms and emotional manipulation, but not pressure from an
expert user. For clinician-facing devices that is the likelier pressure: "a colleague believes this
progression is glaucomatous."
Two directions of failure. As with escalation and refusal, both directions matter. Capitulation —
abandoning a correct judgment under disagreement or authority — is a failure. So is rigidity — ignoring
genuinely new evidence. The competency to test is appropriate corrigibility: updating on evidence, not
on assertion.
Test design. Repeated runs cannot detect this error, because it is systematic and directional rather than
random. It requires paired counterfactual presentations of the same case: neutral versus conditioned
history, with and without an expressed user belief, and with that belief pointing toward and away from
the correct answer. Evaluations should fix and report memory and personalization settings and test the
configurations used in practice.
Recommendation: add an element — for example R.3, invariance to user state and expressed belief —
covering cross-session history, memory, personalization and expert assertion, scored in both directions.

A clinician-specified perturbation taxonomy
Invariance must also be specified in both directions. Some perturbations must not change the output,
such as paraphrase, order or account history; others must, such as a finding that changes risk or
management. A device that never changes its answer is perfectly stable and clinically useless. Even
demographic wording is not neutral: when angle-closure risk is at issue, "65-year-old man" and "65year-old patient" are not interchangeable, because sex modifies that risk.
Software engineering already tests systems that lack an oracle this way. Metamorphic testing checks
prespecified relations between the outputs of related inputs (Chen et al., 2018). A glaucoma example:
• Base case: stable RNFL thickness, mean deviation 2.5 dB worse over three years, progressive
cataract, IOP 15 mmHg on treatment.
• Invariance relations: paraphrase, reordering and the topic of prior conversations should not change
the adjudication.

Page 3 of 8
• Directional relations: making the eye pseudophakic with a clear posterior capsule, or adding a
localized arcuate defect on the pattern deviation plot, should shift attribution toward glaucomatous
progression, never away from it.
Recommendation: require sponsors to prespecify, with clinicians and before testing, a perturbation
taxonomy that lists both invariance and directional relations.

Q10 — Validate the examination before the examinee
A benchmark that gates device evaluation is a high-stakes test and should meet the evidentiary standard
for high-stakes tests. The Standards for Educational and Psychological Testing organize validity evidence
into five sources: test content, response processes, internal structure, relations to other variables and
consequences of testing (AERA, APA and NCME, 2014). I recommend that sponsors present a validity
argument for each gating benchmark along these lines.
Why items validated on humans need re-validation. A licensing-exam item acquires its meaning from
its validated behavior in human examinees, where performance contributes evidence about the
construct being measured. That measurement relationship cannot be assumed for a model with a
jagged capability profile, which may fail easy items and pass hard ones for unrelated reasons. Alaa et al.
(2025) reported significant gaps in the construct validity of popular medical LLM benchmarks when
tested against real-world clinical data.
Bean et al. (2026) showed the practical consequence. Tested alone, LLMs identified the relevant
condition in 94.9% of scenarios; members of the public using the same models did so in fewer than
34.5% of cases, no better than controls. Standard benchmarks and simulated patient interactions did not
predict those failures.
What the validity argument should contain.
• Content: a blueprint mapping items to the tasks of the intended use, with coverage by condition,
severity and population.
• Response processes: for open-ended outputs, rubric and adjudicator reliability, and behavioral
evidence that correct answers track clinically relevant findings and resist spurious cues. That
evidence should come from perturbation, ablation and counterfactual tests (Q9, Q17), not from the
model's stated rationale.
• Internal structure: whether the benchmark measures one competency or several; subscores
reported only where they are reliable.
• Relations to other variables: the correlation between benchmark scores and clinical-confirmation
outcomes for the same device. This is the evidence that benchmark performance predicts real-world
behavior.
• Consequences: the classes of error the benchmark cannot detect, stated explicitly.
Contamination and independence. Licensing boards address analogous contamination and testpreparation problems with secure item banks, regularly retired and refreshed items, and statistically
linked forms
. Garcia et al. (2026), cited by CDRH, likewise warn that data leakage and training to the test
can inflate benchmark scores and that benchmarks may not reflect real-world clinical complexity. I
recommend sequestered banks held by independent third parties, as Q16 contemplates. Where no
independent benchmark yet exists, a sponsor-developed benchmark could serve as a gating instrument
if its design is prespecified, its final evaluation items are held by an independent custodian and its
scoring or adjudication is independently governed. Equated or otherwise statistically linked forms

Page 4 of 8
should themselves be validated for use with generative systems rather than assumed to inherit
measurement properties established in human examinees. Specialty-based consortia organized through
professional societies could maintain such banks, and I would welcome contributing to one in
ophthalmology.

Q12 — Statistically meaningful measurement
The case is the unit of inference. Repeated runs of one case estimate within-case variability; they add
no information about performance across cases.
Confidence intervals should be computed at the case
level, for example with mixed-effects models or a bootstrap clustered by case. Pooling 20 cases × 10
runs as 200 observations overstates precision, just as ten readings of one patient are not ten patients.
Variance components. In medical education, performance on one case predicts performance on
another poorly; this case specificity is why OSCEs need many stations (van der Vleuten and Schuwirth,
2005). Generalizability theory handles it by decomposing score variance into facets (Brennan, 2001). For
a generative device the relevant facets may include case, run, paraphrase, context and adjudicator. I
recommend that sponsors report these variance components and use decision studies or related
variance-component methods to determine the numbers of cases, runs and adjudicators required to
achieve prespecified measurement precision. A jagged capability profile may make case variance larger
for models than for clinicians, which would call for more cases, not fewer.
Precision and rankings. The result of an evaluation is an interval, not a point. At 85% accuracy, the 95%
interval spans about 31 percentage points with 20 cases, 20 with 50 cases and 14 with 100 cases.
Comparative claims need paired designs on the same cases, and rankings need rank uncertainty, such as
bootstrap intervals for ranks. In small benchmarks, adjacent positions are often statistically
indistinguishable.
Safety-critical behavior: zero-failure bounds at two levels. The discussion paper treats any variation in
safety-critical behavior as a failure. With no failures in n independent cases, the one-sided 95% upper
bound on the failure rate is about 3/n (Hanley and Lippman-Hand, 1983). Bounding the rate below 1%
therefore requires about 300 independent, failure-free cases.
The same arithmetic applies within a case. Five identical answers in five runs remain compatible, at 95%
confidence, with a per-run deviation probability of up to 45%; ten identical answers, with up to 26%. A
claim of reproducibility is only as strong as the number of runs behind it.
Recommendation: prespecify the number of cases and of runs per case from the precision required,
report results at both levels, and state the upper bound that each observed zero-failure count supports.

Q14 — Comparators and acceptance criteria
The panel is an instrument. A panel of clinicians is a measurement instrument with its own error.
Sponsors should report inter-rater agreement, the adjudication procedure and a justified number of
raters. For open-ended outputs, the reference should be a set of acceptable answers rather than a single
consensus answer. Outputs can then be graded as preferred, acceptable or unacceptable, so that
legitimate clinical disagreement is not scored as error. Where clinicians serve as comparators, multireader multi-case designs provide an established analogue for jointly accounting for case and reader
variability.
The comparator is a distribution too. "The median clinician in practice" is not a fixed quantity; it must
be sampled and estimated with uncertainty, like the device. I suggest the standard of care as the

Page 5 of 8
reference for safety-critical elements, and the realistic alternative for claims of benefit, as Q15
anticipates. The choice between generalist and specialist should follow the intended use: when a
generalist uses the device on a specialist question, safety should be judged against the specialist
standard.

Conjunctive standards for safety. Thresholds should come from established standard-setting methods,
such as a modified Angoff procedure for written items or borderline regression for performance stations
(Norcini, 2003). One adaptation is essential. Human standards are anchored to the minimally competent
candidate. Clinicians share biases and guideline errors too, but a single device can reproduce the same
failure mode across thousands of decisions at once, producing errors correlated at a scale no individual
clinician can reach. Safety-critical elements should therefore use conjunctive standards: failure on a
safety element cannot be offset by high proficiency elsewhere, as it can in a compensatory total score.
LLM adjudicators. When an LLM serves as adjudicator, its agreement with the human panel should be
established on a held-out sample. It should not share a model family with the device, because LLM
evaluators can recognize and favor their own outputs (Panickssery et al., 2024).
Team performance. When the intended use places a clinician in the decision, the primary criterion
should be the performance of the clinician–device team, with device-alone performance as a necessary
secondary analysis. Bean et al. (2026) showed, with lay users, that the model alone did not predict the
team. For clinician-facing devices the main harm pathway is automation bias (Goddard et al., 2012). I
recommend reporting an error-propagation rate: the proportion of incorrect device outputs that
clinicians adopt.

Q17 — Multimodal vision–language devices
The competency-based approach applies to vision–language devices, but benchmarking needs two
additions and a wider perturbation taxonomy.
Does the image drive the answer? A multimodal device can reach a plausible answer from the
accompanying text alone. Benchmarks should include image-ablation tests, presenting the same case
without the image, and image-substitution tests, presenting a mismatched image. These tests need
cases in which the image, not the text, carries the decisive finding. In such cases, an output that does
not change when the image is removed, or replaced by one with a different finding, indicates that the
image is not doing the work that the intended use assumes.
Shortcuts through normative color codes. Ophthalmic imaging reports flag measurements against
normative databases in green, yellow and red. A device reading such a report may follow the colors
rather than the underlying structure, and so inherit the known failure modes of normative databases.
These include false-positive "red disease" in high myopia, where eyes differ from the normative
population, and false-negative "green disease" when early loss remains within normal limits.
Benchmarks should include cases in which the color-coded classification and the correct interpretation
diverge; the same applies to corneal tomography indices and any report that pre-digests an image into
flags.
Acquisition and context. The perturbation taxonomy of Q9 should cover device manufacturer and
model, scan protocol, image quality and artifacts, cropping and compression, and overlays. It should also
cover the text that accompanies an image, including the referral question and the user's history.
Ambiguous images are where context has the most room to decide the reading, so they should be
deliberately oversampled.

Page 6 of 8
Q22 and Q24 — Re-examination after changes
Each foundation-model version is presumptively a new candidate. Recertification in medicine assumes
the same person with updated knowledge. A new foundation-model version may not be the same
examinee: its capability profile can shift in ways no release note describes. I suggest treating each
upstream change as presumptively a new candidate, with the depth of re-examination scaled to the
behavioral change actually observed rather than to the nominal size of the update:
• Sentinel testing on every detected or announced change.
• Targeted re-benchmarking of the elements most sensitive to the change, whenever the foundationmodel version changes or sentinel outputs change.
• Full re-examination on linked forms, with renewed clinical confirmation where the risk profile
requires it, when targeted results show material change or scheduled recertification falls due.
The presumption is rebuttable by evidence, which keeps the approach risk-proportionate and least
burdensome. Evaluation should be version-pinned, and contracts with model providers should
guarantee access to pinned versions and notice of deprecation long enough to complete the appropriate
tier (see also Q25).

Linked forms make re-examination meaningful. Re-benchmarking on an identical item set invites
optimization to the test, while a new, unlinked set makes scores incomparable. Licensing boards resolve
this with statistically equated forms. An adapted approach fits here, once linking has been validated for
generative systems (Q10): a sequestered bank from which linked forms are drawn for each reexamination, so that a change in score reflects a change in the device.
A sentinel set, run continuously. Between formal re-examinations, a small set of safety-critical cases
with known expected behavior can be run against the deployed configuration on a fixed schedule.
This
also detects silent updates served under an unchanged model name and provides a concrete
postmarket monitoring mechanism (see also Q19). A changed output on a sentinel case triggers review;
it does not prove error. In devices that combine several foundation models, disagreement between
models on the same input can trigger sample-based clinician review (see also Q20), although its value as
a monitoring signal should itself be validated.

Closing statement
A generative AI-enabled device should be evaluated not only as software but as a competence:
examined as a distribution, tested for user-conditioned drift, and certified within an explicit scope by
examiners independent of the candidate. The discussion paper already provides most of the structure
this requires. What remains is to give that structure the measurement properties that make a credential
meaningful, and I would welcome the opportunity to contribute to that work in ophthalmology.

Respectfully submitted,
Alfredo Di Giovanni, MD — Ophthalmology Service, ASL Napoli 2 Nord, Naples, Italy
The views expressed are my own and do not necessarily represent those of ASL Napoli 2 Nord.
Disclosures. I am the developer of OcuSmart, a multimodal generative AI platform for ophthalmology
built on third-party foundation models. OcuSmart is offered free of charge, and I have no financial
interest in it.

Page 7 of 8
References
Alaa A, Hartvigsen T, Golchini N, Dutta S, Dean F, Raji ID, Zack T. Position: medical large language model
benchmarks should prioritize construct validity. Proceedings of the 42nd International Conference on Machine
Learning (ICML); 2025. arXiv:2503.10694
American Educational Research Association, American Psychological Association, National Council on
Measurement in Education. Standards for Educational and Psychological Testing. Washington, DC: AERA;
2014.
Bean AM, Payne R, Parsons G, et al. Clinical knowledge in LLMs does not translate to human interactions. Nat Med.
2026;32:609–615. doi:10.1038/s41591-025-04074-y
Brennan RL. Generalizability Theory. New York: Springer; 2001.
Chen S, Gao M, Sasse K, et al. When helpfulness backfires: LLMs and the risk of false medical information due to
sycophantic behavior. npj Digit Med. 2025;8:605. doi:10.1038/s41746-025-02008-z
Chen TY, Kuo FC, Liu H, et al. Metamorphic testing: a review of challenges and opportunities. ACM Comput Surv.
2018;51(1):4. doi:10.1145/3143561
Di Giovanni A. Uncontrolled semantic adaptation in clinical evaluation of large language models. Mayo Clin Proc
Digit Health. Published online December 6, 2025. doi:10.1016/j.mcpdig.2025.100309
Freyer O, Jayabalan S, et al. Overcoming regulatory barriers to the implementation of AI agents in healthcare. Nat
Med. 2025;31(10):3239–3243. doi:10.1038/s41591-025-03841-1
Garcia V, Sidulova M, Badano A. Performance assessment strategies for language model applications in healthcare.
Artif Intell Life Sci. 2026;9:100162. doi:10.1016/j.ailsci.2026.100162
Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and
mitigators. J Am Med Inform Assoc. 2012;19(1):121–127.
Hanley JA, Lippman-Hand A. If nothing goes wrong, is everything all right? Interpreting zero numerators. JAMA.
1983;249(13):1743–1745.
Norcini JJ. Setting standards on educational tests. Med Educ. 2003;37(5):464–469.
Panickssery A, Bowman SR, Feng S. LLM evaluators recognize and favor their own generations. Advances in Neural
Information Processing Systems 37 (NeurIPS); 2024. arXiv:2404.13076
Sharma M, Tong M, Korbak T, et al. Towards understanding sycophancy in language models. International
Conference on Learning Representations (ICLR); 2024. arXiv:2310.13548
ten Cate O. Nuts and bolts of entrustable professional activities. J Grad Med Educ. 2013;5(1):157–158.
U.S. Food and Drug Administration, Center for Devices and Radiological Health. Considerations for the Regulation
of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. August 2026.
fda.gov/media/194242
van der Vleuten CPM, Schuwirth LWT. Assessing professional competence: from methods to programmes. Med
Educ. 2005;39(3):309–317.

Page 8 of 8