The FDA Shows More of Its Hand on Clinical GenAI

A closer read of the FDA’s August discussion paper reveals an emerging lifecycle framework for regulated Clinical GenAI.

By Scott Walchek, CEO & Co-Founder
Abstract

The FDA’s August 2026 discussion paper on generative AI-enabled medical devices offers its clearest public view yet of an emerging regulatory framework for Clinical GenAI. Read across the paper, the Agency appears to be organizing the category around three lifecycle questions: how clinical risk should be assessed, how clinical competence should be demonstrated, and how performance should be assured as the technology evolves. This Dispatch examines that framework through the lens of RecovryAI’s more than two years of direct FDA engagement developing patient-facing clinical AI. It also sets out our positions on several issues likely to shape the category, including the boundaries of clinical autonomy, physician ground truth as the clinical reference standard, simulation of rare and high-consequence presentations, and independent post-market performance verification.

Earlier this month, the FDA published Considerations for the Regulation of Generative AI-Enabled Medical Devices, its most detailed public discussion yet of how generative AI may fit within medical-device regulation.

It is a discussion paper and request for feedback, not guidance. Read as a whole, though, it begins to provide a clearer structure for a set of regulatory challenges that until now have often been discussed separately.

I read the paper as organizing the landscape around three lifecycle questions: How should clinical risk be assessed? How should clinical competence be demonstrated? And how should clinical performance be assured over time?

The first looks at what the AI does, how independently it does it, and the consequence if it is wrong. The second addresses how an open-ended system can demonstrate competence when exhaustive input-output testing is impractical. The third extends that evidence beyond authorization as models, software and clinical inputs change.

Click to enlarge

The paper asks stakeholders to respond to 26 questions. Twenty-four fall within those three core lifecycle areas: six on risk assessment, eleven on premarket evaluation, and seven on post-market monitoring and change control. Two additional questions address foundation-model device master files and agentic AI.

The question count is not a formal weighting, but it provides another way to read the paper. Nearly half of the core questions concern premarket evaluation, with substantial attention also devoted to what happens after a device enters use.

For RecovryAI, much of this terrain is familiar. By our count, roughly 18 of the FDA’s 24 core lifecycle questions touch challenges we have been building for, testing, or working through directly with the Agency for more than two years.

That experience has shaped several convictions about how this category should be regulated.

1. The regulatory lane should follow the clinical function

The FDA’s risk discussion gets directly into one of the central boundaries in Clinical GenAI: when information becomes clinical direction.

The paper describes a continuum from general information through increasingly action-directing output. It also observes that adding language such as “talk to your doctor” or “I am not a medical professional” does not necessarily make an otherwise action-directing output less directive. The substance and context of what the AI tells the patient remain part of the analysis.

The FDA then adds a second dimension. Risk depends not only on how independently the AI directs or takes action, but also on the potential harm if the output is wrong. Its proposed two-axis heuristic places increasing independence of device activity on one axis and increasing consequence of error on the other.

The FDA’s Two-Axis Risk Framework
Severe Moderate Limited CONSEQUENCES Informational: Non-Directive Informational: Action-Directing Action-Taking: HCP-Supervised Action-Taking: Fully Autonomous ACTIVITY INCREASING RISK
Redrawn from FDA CDRH, Considerations for the Regulation of Generative AI-Enabled Medical Devices, discussion paper, August 2026, Figure 1. A discussion paper, not guidance.

We see a meaningful distinction between AI that provides general health information, AI that supports the judgment of a licensed clinician, and patient-facing AI that independently exercises medical judgment. (I wrote about these distinctions in my April Dispatch “The Emerging Third Lane of Healthcare AI”.)

As a system moves toward greater clinical independence, its intended use, clinical boundaries, handling of uncertainty, escalation behavior, clinician visibility and auditability increasingly become part of the regulatory analysis.

For us, clinical autonomy does not imply unrestricted AI practicing broadly across medicine. It can describe carefully bounded independence within a defined clinical role, supported by evidence appropriate to that role. However, we firmly hold that some degree of clinical autonomy is the key to decoupling healthcare’s supply constraint from human labor. (I wrote about this solution in my July Dispatch “Two Economists Make the Argument for Virtual Care”.)

2. Physician ground truth should remain the clinical reference point

Open-ended conversational AI presents a different validation challenge from conventional software. The possible combinations of language, literacy, context, symptom complexity, and patient presentation make exhaustive testing unrealistic.

The FDA is exploring a competency-based approach influenced, at a high level, by the way clinicians are evaluated. Under that concept, the complete user-facing device would demonstrate defined competencies through benchmarking, followed by clinical confirmation appropriate to its intended use and risk.

The paper considers several forms of clinical confirmation, including requiring real patient inputs, clinician adjudication, standardized patient interactions and prospective studies. It also contemplates independent multi-clinician adjudication as a way to establish the clinical reference against which device output can be compared.

Our position is that physician ground truth should remain the gold standard for clinical validation.

For patient-facing AI exercising medical judgment, performance should ultimately be measured against independent physician judgment. Prospective evidence can establish performance in real clinical use, while benchmarking, simulation and other scalable methods can broaden the range of circumstances evaluated without replacing the human clinical reference standard.

3. Rare, high-consequence presentations need another testing method

A prospective study can produce extensive evidence of routine clinical performance while still encountering relatively few examples of an uncommon complication. Increasing enrollment indefinitely is unlikely to be a practical way to assemble sufficient numbers of every rare, high-consequence presentation.

The FDA’s discussion of competency testing contemplates synthetic data generation and simulation tools, including virtual patient approaches, as potential ways to extend evaluation. It also considers synthetically generated inputs as a supplement where real patient inputs are limited.

Our Hybrid Conversation Generation framework was developed for this purpose. Carefully constructed simulated patient engagements allow us to deliberately test uncommon, difficult and high-consequence situations beyond what a prospective trial may practically accumulate.

The simulation expands the range of presentations tested. However, every clinical judgment produced through those engagements is still evaluated against independent physician ground truth.

4. Performance assurance has to extend beyond authorization

The FDA devotes seven of its core questions to post-market monitoring and change control, including re-benchmarking, clinician review, performance degradation and changes to underlying technology.

A GenAI-enabled medical device can change in ways traditional software may not. Foundation models can be updated by third parties. Medical reasoning frameworks and software can be refined. The clinical information available to the system can expand as wearables and other connected technologies contribute new physiologic signals.

Our view is that a regulated Clinical GenAI device should have a verification capability outside the clinical reasoning system itself that can independently assess whether clinical performance remains consistent as those elements change.

That is the role our Automated Reference Standard, or ARS, is being built to perform: a separate deterministic layer for longitudinal surveillance and reproducible assessment against clinician-established reference standards.

Changes with the potential to affect clinical performance should also trigger defined re-evaluation. Where appropriate, a Predetermined Change Control Plan can establish with the FDA in advance which categories of future changes may be made and how they will be assessed.

Why the FDA’s questions are familiar to us

Our first informational meeting with the FDA on patient-facing clinical AI was in August 2024. Since then, nine formal pre-submission packages have taken us through clinical reference standards, study design, model controls, rare-event evaluation and lifecycle performance monitoring. One consolidated submission alone mapped 68 separate items of Agency feedback into product, validation and regulatory responses.

That history explains our familiarity with much of the paper. The FDA is now inviting the broader industry into a public discussion around questions we have been working through directly with the Agency as part of building and validating the core components upon which our VCA is constructed.

What the publication adds

The paper gives the industry a much clearer public reference point for the regulatory landscape around Clinical GenAI.

For RecovryAI, our regulatory work continues toward demonstrating the safety and effectiveness of the VCA for its defined intended use. Many of the issues raised in the paper are already reflected in our architecture, validation program and lifecycle controls.

Our larger thesis is that patient-facing clinical AI, once authorized to exercise medical judgment independent of a licensed provider, can create clinical capacity that is no longer constrained by clinician labor. That is an entirely different proposition from using AI primarily to help the existing workforce process clinical work more efficiently.

Clinical AI that independently exercises medical judgment with patients should earn the confidence placed in it. We have built around that premise from the beginning, and through both the FDA’s public comment process and our ongoing regulatory work, we intend to contribute practical evidence and experience to the developing regulatory framework for this category.

Regulatory status: The Virtual Care Assistant is an investigational device, limited to investigational use under U.S. law. It has received Breakthrough Device Designation from the FDA and is not authorized for commercial distribution. Nothing in this Dispatch predicts the outcome of any regulatory submission.

Latest

Stay Connected