Considerations for the Regulation of Generative AI-Enabled Medical Devices · FDA discussion paper · August 2026
A synopsis prepared by recovry.ai
Increasing risk ↗

The future ofhealthcare AIregulation

is being shaped by the FDA just asked.
InformsDirectsActs with oversightActs alone

In August 2026 the FDA published its first full account of how it might regulate generative AI in medicine: a discussion paper that proposes no rules, asks 26 questions, and invites anyone to answer them. What follows summarizes its three areas of focus and the questions that matter most within each, because the answers will lay the groundwork for what generative AI is allowed to do in care, and how it is held to account.

The proposed risk framework · Figure 1 redrawn · Questions 1 to 6

GenAI risk on two axes: action and consequence.

The FDA’s proposed framework places every GenAI function on two axes: what the device does, and the consequence of relying on an incorrect output. Hover or tap a square to see the FDA’s examples for that spot.

Consequences
Increasing risk ↗
SevereModerateLimited
Informational:
Non-Directiveinforms
Informational:
Action-Directingdirects
Action-Taking:
HCP-Supervisedacts with oversight
Action-Taking:
Fully Autonomousacts alone
Activity →
Hover a square
Nine worked examples sit across this map.
From a cardiovascular risk score in the lower-left to an autonomous stroke order set in the upper-right, each quoted from Section IV.
Paper section IV · read in the paper

Assess clinical risk

What happens if the AI is wrong?

Six questions on how the risk of a GenAI function should be assessed, plus a proposed map that places each one by what it does and what a wrong output could cost.

Click any question for its full wording from the paper, and further analysis
The directiveness dial · Question 2

The continuum from non‑directive to action‑directing.

General information about a drug informs. A specific dosing instruction directs. The FDA’s proposed framework treats everything in between as degrees, and says directiveness depends on the substance and context of an output, not on words like “recommend.” A “talk to your doctor” line may not make it any less directive.

Drag the dial through four phrasings of the same advice, quoted from Section IV.

Drag to move from informing to directing
Non-directiveAction-directing
Paper section V · read in the paper

Demonstrate clinical competence

How is the AI shown to be safe and effective?

Eleven questions on evidence. You cannot test every input to an open-ended system, so the FDA proposes a competency-based approach modeled on how clinicians are credentialed: benchmark first, then confirm in clinical use.

Click any question for its full wording from the paper, and further analysis
The competency-based approach · Section V · Questions 7 to 17

Two steps to demonstrate clinical competence: benchmark the device, then confirm it in clinical use.

Exhaustive testing of every input is impractical for open-ended systems, so the proposed approach borrows from how clinicians are credentialed: structured assessment first, then supervised real-world practice. The final, deployed device is what gets evaluated, not the foundation model on its own.

Ten competencies to benchmark
First the device is tested without patients, across ten named competencies. Select one to see what it probes.
Five ways to confirm clinical competence
Then it is confirmed in real or clinically representative use, by one or more of these five methods, listed from least to most rigorous. Select a method.
One caveat, verbatim: clinical confirmation “might not require a prospective clinical study in every case.” That is Question 11.
Paper section VI · read in the paper

Assure clinical performance

What happens after the AI is authorized?

Seven questions on life after authorization, plus two on foundation models and agentic AI. The central trade is stated out loud: more uncertainty before market, offset by heavier monitoring after.

Click any question for its full wording from the paper, and further analysis
After authorization · Section VI · Questions 18 to 24

Two ways to assure post-market performance: monitor the AI in use, and re-test it when it changes.

Under the proposal, the benchmark a device passed before authorization would become its baseline. Monitoring would watch for drift from that baseline, and a change to the device, or to the model beneath it, may trigger a re-test against it.

Three proposed ways to monitor in use
Question 20 also asks whether AI supervisory agents could help with this.
Three kinds of change that may trigger a re-test
The third is the hard case: the model’s developer, not the device maker, changes it (Question 24).
All twenty-six · click any question for the full verbatim wording from the paper, and further analysis · filter by who it touches

All 26 questions.

Verbatim