In August 2026 the FDA published its first full account of how it might regulate generative AI in medicine: a discussion paper that proposes no rules, asks 26 questions, and invites anyone to answer them. What follows summarizes its three areas of focus and the questions that matter most within each, because the answers will lay the groundwork for what generative AI is allowed to do in care, and how it is held to account.
The FDA’s proposed framework places every GenAI function on two axes: what the device does, and the consequence of relying on an incorrect output. Hover or tap a square to see the FDA’s examples for that spot.
Six questions on how the risk of a GenAI function should be assessed, plus a proposed map that places each one by what it does and what a wrong output could cost.
General information about a drug informs. A specific dosing instruction directs. The FDA’s proposed framework treats everything in between as degrees, and says directiveness depends on the substance and context of an output, not on words like “recommend.” A “talk to your doctor” line may not make it any less directive.
Drag the dial through four phrasings of the same advice, quoted from Section IV.
Eleven questions on evidence. You cannot test every input to an open-ended system, so the FDA proposes a competency-based approach modeled on how clinicians are credentialed: benchmark first, then confirm in clinical use.
Exhaustive testing of every input is impractical for open-ended systems, so the proposed approach borrows from how clinicians are credentialed: structured assessment first, then supervised real-world practice. The final, deployed device is what gets evaluated, not the foundation model on its own.
Seven questions on life after authorization, plus two on foundation models and agentic AI. The central trade is stated out loud: more uncertainty before market, offset by heavier monitoring after.
Under the proposal, the benchmark a device passed before authorization would become its baseline. Monitoring would watch for drift from that baseline, and a change to the device, or to the model beneath it, may trigger a re-test against it.