← All 95 filings

Amr Saad, MD (Pallas Kliniken)

CliniciansClinicianFiled September 7, 20261,608 words · 1 attachmentFDA-2026-N-7874-0070
“The review step named in the labeling therefore carries a responsibility the intended use has already made impossible to exercise.”

What they argued

RecovryAI’s one-line reading of the filing.

Q14 only: human-AI team performance is basis whenever accountable clinician sees output, but workflow position must be fixed in intended use; IDx-DR example.

Themes it raises

3 of the 21 themes in the docket, each with the passage we counted, verbatim.
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“Human-AI team performance should be the basis for evaluation whenever a clinician who remains accountable for the resulting decision will see the device output.”
Who is accountable when something goes wrongFDA Q21
“A third question follows from those and is the one the IDx-DR labeling leaves open, namely which identified person is accountable for the output in a workflow where no clinician assessment precedes it.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“The review step named in the labeling therefore carries a responsibility the intended use has already made impossible to exercise.”

FDA questions it names

Questions this filing names by number.

Q14 · Comparators and acceptance criteria

Coded positions

Where a position was recorded question by question.
Q14For open-ended AI outputs, who is the performance comparator: a clinician panel, generalists, specialists, or the human-AI team?
Evaluate the clinician and AI working together

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
No position stated
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
No position stated
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
No position stated
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Not stated
High-consequence work: Not stated
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

This comment responds to one sentence of Discussion Question 14: "When should human-AI team performance, rather than the device operating alone, serve as the basis for evaluation?" It does not address the other questions. The full comment is attached as a PDF.

Summary. Human-AI team performance should be the basis for evaluation whenever a clinician who remains accountable for the resulting decision will see the device output. Team performance cannot be measured, however, unless the position of the output in the workflow is fixed as part of what is being evaluated. Section V.D.1 notes that comparator selection "might include how the device is actually used, including the combined performance of the clinician and device working together as a human-AI team versus the device working in a fully autonomous workflow." That phrase, "how the device is actually used," is underspecified in one respect that has already produced a concrete problem in an authorized device class.

In the De Novo summary for IDx-DR, CDRH states that the clinical study "demonstrated safe and effective clinical performance of IDx-DR when used to automatically (without physician assistance) detect mtmDR" (DEN180001, page 12). The labeling in the same document nevertheless states that "Physicians should review IDx-DR results and advise patients of recommended referrals to an eye care provider" (DEN180001, page 2). What reaches that physician is a classification, not the images. In the pivotal trial the system "correctly identified 173 of the 198 fully analyzable participants with fundus mtmDR," an observed sensitivity of 87.4 percent, and the remaining 25 are known only because a reading center graded every image for study purposes. In routine use, by design, no one does. The review step named in the labeling therefore carries a responsibility the intended use has already made impossible to exercise.

The sequence also matters to patients and can be measured. In a preregistered vignette experiment with 570 members of the general public (Schaffernak et al., J Med Internet Res 2026;28:e93172; I am a coauthor), estimated marginal means for trust in the medical decision were 5.17 with no AI, 4.87 when the physician reviewed the case and a diagnostic AI suggestion concurrently, and 5.35 when the physician assessed the case independently first. Concurrent diagnostic AI was rated significantly lower than the no-AI baseline (t565=2.65; P=.04) and than sequential diagnostic AI (t565=4.13; P<.001). The device, the output and the clinician were held constant. Only the order changed.

Recommendation. For a GenAI-enabled device whose output is read by a clinician who retains responsibility for the decision, the workflow position of that output should be treated as part of the device under evaluation rather than as a deployment variable, and specified in the intended use with the same precision as the output itself. Two things determine whether a human-AI team exists at all: the point at which the output is displayed to the clinician, and whether an independent clinical assessment is recorded before it is displayed. Until those are fixed, the same device and the same clinicians will yield different team performance under different orderings, and a sponsor could reasonably test the ordering that performs best while sites deploy the one that is fastest.

Amr Saad, MD MBA
Pallas Kliniken, Olten, Switzerland
ORCID 0000-0002-0574-6739

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Comment on Docket FDA-2026-N-7874, Document FDA-2026-N-7874-0001, "Considerations
for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request
for Feedback" (posted August 18, 2026).

This comment addresses one sentence of Discussion Question 14 and nothing else in the paper. The sentence is:
"When should human-AI team performance, rather than the device operating alone, serve as the basis for
evaluation?" I take the paper at its own description of its status. It "is intended for discussion purposes only and
does not represent draft or final guidance," and it "is not intended to propose or implement policy changes
regarding how CDRH intends to regulate generative AI-enabled devices." Nothing below is offered as a response
to a proposed requirement.
My answer is that human-AI team performance should be the basis for evaluation whenever a clinician who
remains accountable for the resulting decision will see the device output
, and that team performance cannot be
measured at all unless the position of the output in the workflow is fixed as part of what is being evaluated.
Section V.D.1 already approaches this. It notes that considerations for selecting the comparator "might include
how the device is actually used, including the combined performance of the clinician and device working
together as a human-AI team versus the device working in a fully autonomous workflow, depending on the
intended use." My submission is that "how the device is actually used" is underspecified in one respect that has
already produced a concrete problem in an authorized device class, and that generative outputs will reproduce it
at a much larger scale.
Autonomous diabetic retinopathy screening is the case in which the division of labor between device and
clinician was settled by an authorization rather than by local clinical practice. In the De Novo summary for IDxDR, CDRH states that the clinical study "demonstrated safe and effective clinical performance of IDx-DR when
used to automatically (without physician assistance) detect mtmDR" (DEN180001, page 12; De Novo request
received January 12, 2018, granted April 11, 2018). The pivotal trial describes the system in the same terms, as
autonomous in the sense of "without human expert reading of the retinal images," and states that responsible
implementation in primary care "requires autonomy (i.e., a use case that removes the requirement for review by
human experts)" (Abràmoff et al., npj Digital Medicine 2018;1:39). The labeling reproduced in the same De
Novo summary nevertheless assigns a task to a professional: "Physicians should review IDx-DR results and
advise patients of recommended referrals to an eye care provider for evaluation and potential treatment"
(DEN180001, page 2).
Those two statements are compatible only if reviewing a result means something other than reviewing a finding.
What reaches the physician is a classification, not the images. In the pivotal trial the system "correctly identified
173 of the 198 fully analyzable participants with fundus mtmDR," an observed sensitivity of 87.4 percent
(173/198). The 25 participants in that trial who had referable disease and did not receive a positive output are
known only because a reading center graded every image for study purposes. In routine use, by design, no one
does. A physician who is asked to review the result and advise the patient has no means of separating that group
from true negatives, so the review step named in the labeling carries a responsibility that the intended use has
already made impossible to exercise. This is not a failure mode of the device. It follows from where the device
was placed in the sequence, and that placement was part of what was authorized.
The sequence also matters to patients, and it can be measured. Two preregistered vignette experiments with 489
and 570 members of the general public in Germany varied whether a physician used no AI support, descriptive
AI support, or diagnostic AI support, and in the second study varied the timing of that support (Schaffernak et
al., Journal of Medical Internet Research 2026;28:e93172). In study 2 (N=570, 7-point scales), estimated
marginal means for trust in the medical decisions were 5.17 (95% CI 5.02 to 5.33) with no AI, 4.87 (4.70 to
5.04) when the physician reviewed the case and a diagnostic AI suggestion concurrently, and 5.35 (5.19 to 5.51)
when the physician assessed the case independently first and reviewed the AI output afterward. Concurrent
diagnostic AI was rated significantly lower than the no-AI baseline (t565=2.65; P=.04) and significantly lower
than sequential diagnostic AI (t565=4.13; P<.001). The device, the output, and the clinician were held constant.
Only the order changed.
What I would ask CDRH to consider is the following. For a GenAI-enabled device whose output is read by a
clinician who retains responsibility for the decision, the workflow position of that output should be treated as
part of the device under evaluation rather than as a deployment variable, and specified in the intended use with
the same precision as the output itself. Two things determine whether a human-AI team exists in the first place:
the point at which the output is displayed to the clinician, and whether an independent clinical assessment is
recorded before it is displayed. A third question follows from those and is the one the IDx-DR labeling leaves
open, namely which identified person is accountable for the output in a workflow where no clinician assessment
precedes it.
Until those are fixed, human-AI team performance is not a well-defined quantity, because the same
device and the same clinicians will yield different team performance under different orderings, and a sponsor
could reasonably test the ordering that performs best while sites deploy the one that is fastest. Where team
performance is proposed as the evaluation basis, it would be worth asking which ordering was tested, and
whether that ordering is enforced by the device or merely recommended in labeling.
I have no financial interest in any device named in this comment.

Respectfully submitted,
Amr Saad, MD MBA
Pallas Kliniken, Olten, Switzerland
ORCID 0000-0002-0574-6739

References. FDA/CDRH, "Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and
Request for Feedback," https://www.fda.gov/media/194242/download. FDA, De Novo Summary DEN180001 (IDx-DR),
https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf. Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC.
Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj
Digital Medicine 2018;1:39 (published August 28, 2018). https://doi.org/10.1038/s41746-018-0040-6. Schaffernak I, Cecil J,
Kokje E, Kleine AK, Saad A, Zemo F, Lermer E. Effects of Type and Timing of Clinician-Facing AI Support on Patient Trust in
Medical Consultations: 2 Vignette Experiments. J Med Internet Res 2026;28:e93172. https://doi.org/10.2196/93172.