← All 95 filings

Ben Locwin

IndustryConsultantFiled August 26, 20261,943 words · 1 attachmentFDA-2026-N-7874-0043
“Patient empowerment is a terrible goal, because it implies that the patient would have the ability to differentially-apply the therapy based on their idiosyncratic opinions.”

What they argued

RecovryAI’s one-line reading of the filing.

'Patient empowerment is a terrible goal'; all outputs 'for information only', humans in loop; Q7 yes but comparators from meta-analyses, agentic benchmarking inadequate, third parties inadvisable.

Themes it raises

8 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Layering an ability to detect, such as with an FMEA (Detectability) can help, but traceability is hard (read: impossible) to implement, since the models are opaque.”
Whether the user can judge the outputFDA Q3, Q4
“Patients are not trained in healthcare, usability of devices, or translating and inferring context from generative AI outputs.”
Escalating too little and too muchFDA Q6
“The correct answer here is that Type I errors can be more dangerous than Type II errors.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“There will always be edge cases and corner cases, but this will cover the bolus of submissions.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Giving sponsors a naïve data framework to start with, instead of having them build it internally in a bespoke manner will guide them on how to prevent issues like contamination, saturation, overfitting, etc.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“The synthetic distributions would be the control data, and need to be compared with the real-world distributions, where clinical significance and practical significance are shown to be meaningfully different from each other to be considered for clearance or approval.”
Devices that plan and take actionsFDA Q26
“But the final element of the benchmarking on A.1. Agentic AI Capabilities is woefully inadequate.”
What the rules cost sponsors and the marketNot asked by the FDA
“Bringing a third-party requirement also opens the door to third-party organizations gaming the system and creating a profit center on drawing out these tests.”

FDA questions it names

Questions this filing names by number.

Q7 · The competency-based approachQ9 · The benchmarking structureQ11 · Clinical confirmation without a prospective trialQ15 · Performance against usual care

Coded positions

Where a position was recorded question by question.
Q9Do the ten benchmark competencies, from clinical knowledge to generalizability, add up to enough evidence of safety and effectiveness?
Use the structure, with additions or changes

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Opposes
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
No position stated
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Informs
High-consequence work: Informs
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Response to the framework document:
1.Does the two-axis risk framework, organized around device activity and the consequence of relying on an incorrect output, appropriately capture the dimensions most relevant to the risk of a GenAI-enabled software function? If there are additional dimensions—such as the reversibility of a resulting action, the availability of downstream safeguards, the time pressure of the deployment setting, or the traceability of the output (i.e., to primary source materials)—that should be represented in a risk framework, please describe and provide examples of how they should be represented.
The bi-axial framework can function properly for this application, similarly with the typical risk calculus of R = Pf x Mc (where Pf is the probability of failure, and Mc is the magnitude of consequence). However, that classic equation has also lent itself to uncountable issues in the industry with ‘how’ to measure each properly (qualitatively, subjectively, combined with objective data), including via ICH Q9 and ISO 14971. Layering an ability to detect, such as with an FMEA (Detectability) can help, but traceability is hard (read: impossible) to implement, since the GenAI models are opaque. Humans in- or on- the loop need to be included in the framework for decision-making.
See attached document for responses to the other 16 questions from the Generative AI document.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Discussion questions
Does the two-axis risk framework, organized around device activity and the consequence of relying on an incorrect output, appropriately capture the dimensions most relevant to the risk of a GenAI-enabled software function? If there are additional dimensions—such as the reversibility of a resulting action, the availability of downstream safeguards, the time pressure of the deployment setting, or the traceability of the output (i.e., to primary source materials)—that should be represented in a risk framework, please describe and provide examples of how they should be represented.
The bi-axial framework can function properly for this application, similarly with the typical risk calculus of R = Pf x Mc (where Pf is the probability of failure, and Mc is the magnitude of consequence). However, that classic equation has also lent itself to uncountable issues in the industry with ‘how’ to measure each properly (qualitatively, subjectively, combined with objective data). Layering an ability to detect, such as with an FMEA (Detectability) can help, but traceability is hard (read: impossible) to implement, since the models are opaque. Humans in or on the loop need to be included in the framework for decision-making.
CDRH seeks input on how to account for the spectrum of GenAI-enabled informational functions that vary in the degree to which they direct a user to a particular action (i.e., between “non-directive” and “action-directing”). What characteristics of an output—such as its wording, specificity, personalization, or context—could be considered as modifiers of the risk associated with an informational function, after accounting for the device’s overall functionality and intended use? What additional information would provide manufacturers with sufficient clarity and predictability around risk assessment for informational functions while recognizing that directiveness may exist along a continuum rather than as a binary distinction?
GenAI-enabled functions that are directive are giving medical advice, so all outputs (if not standardized and approved by FDA) need to be ‘for information only’ and have linguistic modifiers required to clarify this.
CDRH seeks input on whether and when a GenAI-enabled function that results in delivery of clinical information to patients, as opposed to HCPs, could present different or higher risks, while also recognizing the potential benefits associated with improved patient empowerment, engagement, and access to clinical information. What device characteristics, output features, or safeguards might mitigate risks that could arise when a user lacks the domain knowledge to independently evaluate an output, without unnecessarily underestimating patient capability?
Yes, unequivocally. Patients are not trained in healthcare, usability of devices, or translating and inferring context from generative AI outputs. Patient empowerment is a terrible goal, because it implies that the patient would have the ability to differentially-apply the therapy based on their idiosyncratic opinions, which is 100% unregulatable.
CDRH seeks input on whether and how the distinction between generalist and specialist physicians might inform the assessment of risk for HCP-facing GenAI-enabled functions. Under what circumstances, if any, might risk be affected when an HCP who lacks the relevant clinical specialist knowledge receives information that falls within a specialist area of practice? What device characteristics or safeguards might mitigate such risks?
The GenAI output needs to include modifiers (as mentioned in #2 above) that explain to generalists that the output is designed for specialists. Otherwise generalists will take the output, impute their opinion and overall (general) license to practice medicine, and go forward with the results.

For multi-turn conversational GenAI-enabled devices that may migrate from providing “non-directive” information to “action-directing” information over the course of an exchange, how should risk be assessed across realistic conversational trajectories? How could the intended use of such a device be characterized when its behavior is emergent across a conversation?
The GenAI needs a model framework that recognizes and self-identifies then it is emergently creating advice and direction, and it is prevented from providing that to the patient with a warning that the conversation has entered that territory.

For GenAI-enabled care escalation functions, CDRH is considering how the evaluation may account for both under-escalation and over-escalation. How could manufacturers characterize and weigh these two directions of error, given that they may not be commensurable and that acceptable trade-offs may vary by clinical context?
The correct answer here is that Type I errors can be more dangerous than Type II errors. Over-escalating can lead to unnecessary treatment harms and diagnostic use. These are errors of commission, rather than under-escalation which would lead to potential errors of omission.
7. Is the competency-based approach described above, i.e., device benchmarking followed by clinical confirmation, a useful and appropriate framework for evaluating GenAI-enabled devices? Yes. There will always be edge cases and corner cases, but this will cover the bolus of submissions.
8. How might the two-axis risk framework described in Section IV be considered within the competency-based approach described above to help determine the level of evidence needed for a premarket submission? See Answer 1 above.
9. Would a benchmarking structure such as the one described above be likely to provide adequate evidence of clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability to support a reasonable assurance of safety and effectiveness? Are there elements that are missing, redundant, or inappropriately categorized? Are there externally developed standards that could be leveraged? The guidelines available in the wild are insufficient to use here. But the final element of the benchmarking on A.1. Agentic AI Capabilities is woefully inadequate. The previous elements are sufficiently robust, but having a single slice on the agentic AI itself doesn’t interrogate ‘what’ it’s doing or ‘how’ it’s doing it deeply enough to prevent risks.
10. CDRH recognizes that publicly available benchmarking assets may be subject to data contamination, saturation, and limited real-world representativeness. How should a sponsor establish that performance on a given benchmark predicts safe and effective real-world behavior for the device’s intended use? What evidence should support the construct validity of a benchmark used to gate device evaluation, and what role should sponsor-developed benchmarks play given potential concerns around independence and optimization to the test? Giving sponsors a naïve data framework to start with, instead of having them build it internally in a bespoke manner will guide them on how to prevent issues like contamination, saturation, overfitting, etc.
11. CDRH is considering that clinical confirmation for a GenAI-enabled device might not require a prospective clinical study in every case, and has described above a range of approaches of increasing rigor and patient exposure. How might a sponsor select and justify a confirmation approach tailored to a device’s intended use and proportionate to the device’s risk profile? Are there device types or risk profiles for which one or more of these approaches would be insufficient or inappropriate? How might the anticipated distribution of real-world inputs be taken into consideration? Are there other methods of clinical confirmation that might help inform the evaluation of GenAI-enabled devices? A risk framework that takes into account therapeutic area/use, patient exposure, and real-world use cases must be implemented.
12. How should sponsors achieve statistically meaningful performance measurement for GenAI-enabled devices? Where synthetically generated inputs supplement real patient data, how should sponsors account for differences between the synthetic and real-world distributions when estimating performance, and under what conditions, if any, is it appropriate to combine benchmarking evidence and clinical confirmation evidence to support a single performance estimate? The synthetic distributions would be the control data, and need to be compared with the real-world distributions, where clinical significance and practical significance are shown to be meaningfully different from each other to be considered for clearance or approval.
13. For which clinical domains, device functions, or subpopulations is synthetic data particularly well-suited, or particularly inadequate, as a supplement to real-world evidence? What safeguards would mitigate the risk that synthetic data generated by models of the same class as the device under evaluation reproduces the very performance gaps the evaluation is intended to detect, particularly for underrepresented subgroups? You could argue that you could create synthetic subpopulations for any non-rare and non-ultra rare condition. But the synthetic data should come from literature metaanalyses, not single studies, to account for large-scale variance.
14. For open-ended device outputs, how should performance comparators and acceptance criteria be selected? When a panel of qualified clinicians serves as the comparator, how should the applicable standard (for example, the standard of care versus the performance of a median clinician in practice) be defined and justified? Should generalist or specialist physicians be used as a performance standard? When should human-AI team performance, rather than the device operating alone, serve as the basis for evaluation? Comparators should be available metaanalytic data for the reasons described above. For performance standards, a mixture of generalists and specialists should be used, because they WILL be used in the real world later on. Treat it like a Gage R&R study, to catch inter-clinician variance.
15. Are there ways in which performance might be assessed relative to the care, technology, or course of action likely to occur in the absence of the device, rather than to the comparators described in this section? What approaches might be used to identify and justify a comparator such as unaided clinical judgment, delayed specialist review, or no intervention? The only reliable comparator is the ‘no intervention’ condition. The others are too unpredictable and variable.
16. Should independent third parties be involved in some or all aspects of a competency-based approach, including device benchmarking and clinical confirmation? If so, in what ways might qualified, independent third-party participation contribute to this approach, and what qualifications and independence criteria should apply? Are there aspects of the assessment for which third-party involvement would be impractical or inadvisable? What safeguards or program design features would be critical to prevent third-party participation from limiting competition or preventing innovation? Bringing a third-party requirement also opens the door to third-party organizations gaming the system and creating a profit center on drawing out these tests. Inadvisable.
17. CDRH is exploring whether a competency-based approach could be applied to devices that incorporate a variety of underlying model architectures, such as multimodal vision-language models and generative or predictive world models. Are there aspects of the described approach that might be ineffective or inapplicable to these devices? What additional or distinct considerations might need to be addressed? Certainly in every case, an assessment of what the architecture includes, such as MMVL models, etc. Those would then be categorized as higher risk, based on A.1 in the framework of your model.