Deborah Ault, RN
“Task completion is not an adequate healthcare-AI safety metric.”
What they argued
Q18 'Sometimes, but only within reasonable limits'; uncertainty shrinks with autonomy/harm; Q7 'Yes, provided competency defined broadly'; 'AI + qualified clinician' safest model.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
Coded positions
Require specialist review or escalation when needed
Set the trade-off for the clinical context
Consider risk factors beyond the two axes
Check for shared blind spots in generated test data
Keep real evidence for claims synthetic data cannot establish
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Specify the tests or controls a future change must pass
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Clarify responsibility and reportable failures
Across the five cross-cutting questions
High-consequence work: Advises
The comment as filed
Please see the attached comments of Deborah “Nurse Deb” Ault, RN, CCM, CCP, AATMC, BCPA, MBA, Founder and President of Nurse Deb Speaking, Media & Consulting, LLC, in response to FDA’s Discussion Paper, “Considerations for the Regulation of Generative AI-Enabled Medical Devices,” Docket No. FDA-2026-N-7874.
These comments focus particularly on patient-facing and agentic AI, clinical safety, evidence-based decision support, risk assessment, escalation, postmarket monitoring, accountability, and the emerging clinical-adjacent AI safety gap.
Attachment
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Response to FDA Discussion Paper: Considerations for the Regulation
of Generative AI-Enabled Medical Devices
Docket No. FDA-2026-N-7874
Cover Letter
To the Food and Drug Administration:
I should begin with an apology.
I missed FDA’s 2025 request for public comment concerning the measurement and evaluation of AIenabled medical-device performance in the real world, which closed on December 1, 2025. In my
defense, my own AI had not yet been properly tasked—or sufficiently trained by its human—to find
federal requests for comment about which I might have something useful to say. [2]
We have since corrected that particular failure mode.
Some of the comments I am submitting now may therefore have fit more neatly within FDA’s earlier
request. I hope you will forgive the late arrival, and I appreciate the opportunity to share those
observations here.
The irony of AI helping me discover that I had previously missed an opportunity to comment on the
safety and oversight of AI is not lost on me.
There is also a substantive lesson in that small example: an AI system can do precisely what a human
asks it to do and still fail the human because the human did not know the right question to ask.
That problem becomes considerably more consequential when the subject is healthcare.
My perspective comes from decades of work at the intersection of patients, clinical care, health
insurance, utilization management, employers, care navigation, healthcare purchasing, and patient
advocacy. I strongly support the responsible use of artificial intelligence in healthcare. Properly
designed, AI has extraordinary potential to expand access to current clinical knowledge, improve
consistency, reduce administrative burden, identify risk sooner, and help clinicians and patients make
better decisions.
But healthcare AI should not merely become better at doing what we tell it to do.
It must become better at recognizing when what we told it to do is not actually what the patient
needs.
This distinction is fundamental because AI can be extraordinarily smart without understanding
human behavior.
Humans frequently do not ask the question they actually need answered.
We ask proxy questions. We minimize symptoms. We avoid frightening possibilities. We worry about
money. We procrastinate. We rationalize. We become embarrassed. We omit information because we
do not know it matters. We describe symptoms incorrectly. We tell ourselves that something is probably
nothing. Sometimes we ask an administrative question because asking the clinical question would
require acknowledging that something might actually be wrong.
A healthcare AI can possess extraordinary medical knowledge and still fail a patient if it does not
understand that fundamental feature of human behavior.
That distinction underlies many of my comments below.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 1
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
I am particularly concerned about the rapidly developing category of patient-facing and agentic AI that
may appear administrative, informational, or operational but can foreseeably influence whether, when,
where, from whom, or what healthcare a person receives. Some of these technologies may fall within
FDA’s medical-device jurisdiction; others may not. From the patient’s perspective, however, the
regulatory classification is considerably less important than the consequence.
A regulatory boundary should not become a patient-safety gap.
My comments therefore focus primarily on risk assessment, patient-facing AI, escalation, competency
evaluation, evidence-based decision support, postmarket accountability, agentic systems, and the
emerging evidence from litigation and state regulation that these risks are no longer theoretical.
The question headings below paraphrase FDA’s consolidated discussion questions; the full wording
appears in Appendix B of the August 2026 discussion paper. [1]
I. Risk Assessment
Question 1 — Is the proposed two-axis risk framework sufficient?
FDA asks whether assessing both the function performed by a generative-AI medical device and the
consequence of relying upon an incorrect output adequately captures risk, and what additional
considerations may be necessary.
Response
The proposed framework is a useful starting point, but I do not believe that “incorrect output” captures
the full universe of foreseeable patient harm.
Healthcare AI can produce an entirely correct output and still create an unsafe interaction.
Consider a patient who asks a health-benefits or navigation AI:
“Can you send me my insurance card?”
The AI retrieves the correct insurance card immediately. The task is completed flawlessly.
But imagine that the patient is a stressed, obese, middle-aged executive with diabetes sitting alone in
his home office late at night. He is sweating, nauseated, and experiencing what he describes to himself
as severe heartburn. He is considering seeking care and is trying to find his insurance information first.
He never asks the AI:
“Am I having a heart attack?”
The AI answers the question he asked.
He returns to work.
At 5:00 a.m., his spouse realizes he never came to bed and finds him dead in his home office.
In that scenario, the AI did not hallucinate. It did not give incorrect clinical advice. It did not malfunction.
It did exactly what it was asked to do.
That is the safety problem.
And it illustrates a principle that I believe deserves explicit consideration in the regulation and evaluation
of patient-facing AI:
AI can be smart without understanding human behavior.
A human being may not ask, “Am I having a heart attack?”
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 2
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
He may ask for his insurance card.
A patient worried about medication toxicity may ask, “How much does this prescription cost?”
A diabetic patient with a potentially limb-threatening foot problem may ask, “Can you find me a
podiatrist?”
A patient whose symptoms have worsened may ask, “Can I move my appointment to next month?”
The literal request may be administrative while the underlying need is clinical.
Risk assessment therefore should include not only the consequence of an incorrect output, but also:
the consequence of an incomplete interaction;
failure to identify a latent clinical need;
failure to recognize safety-relevant context;
risk created by omission;
risk of inappropriate reassurance;
risk that administrative or financial information causes delay or abandonment of necessary care;
whether the system has an opportunity and responsibility to ask an appropriate follow-up question;
whether qualified clinical escalation is available;
the time sensitivity and reversibility of the potential harm; and
the cumulative effect of a multi-turn interaction.
This concern is not hypothetical.
Amazon Science’s 2026 PatientAgentBench evaluated 10 models across four families on 1,200 patientfacing scenarios and found an important divergence between successful task execution and safe
clinical behavior. Triage quality was the most discriminating dimension, and the authors reported that
agents often acted on administrative requests without clinical screening. [3][4]
FDA should review that work as it considers how patient-facing and agentic AI should be evaluated.
PatientAgentBench paper: https://arxiv.org/abs/2607.25485
Amazon Science reference implementation: https://github.com/amazon-science/PatientAgentBench
The significance of this finding extends beyond the particular agents evaluated.
Traditional technology evaluation tends to ask:
Did the system complete the requested task correctly?
Healthcare safety requires an additional question:
Should the system have completed that task without first recognizing that something more
important might be happening?
PatientAgentBench provides empirical support for incorporating this distinction into FDA’s risk
framework.
The appropriate regulatory concept is therefore broader than output accuracy.
Task completion is not an adequate healthcare-AI safety metric.
For patient-facing systems in particular, FDA should evaluate whether the AI can recognize when the
patient’s literal request may be only a proxy for an underlying healthcare need.
Question 2 — How should FDA distinguish informational AI from action-directing AI?
Response
FDA correctly recognizes that “informational” and “action-directing” are not cleanly separated
categories.
In healthcare, information itself can direct action.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 3
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
A patient does not need to receive the grammatical command “Do not seek care” for an AI interaction to
produce that result.
For example:
a cost estimate may cause a patient not to fill a prescribed medication;
an explanation of insurance coverage may cause a patient to postpone treatment;
a provider-ranking system may determine where a patient seeks care;
a statement that a service is “not covered” may function practically as a denial of access;
an AI-generated assessment of urgency may alter whether a patient seeks emergency, urgent,
routine, or no care.
For that reason, regulatory risk should be based partly upon foreseeable influence on patient behavior,
not merely upon whether the system uses imperative language or explicitly recommends an action.
I suggest the following principle:
If technology can foreseeably influence whether, when, where, from whom, or what healthcare a
person receives, the degree of safety oversight and accountability should reflect the degree of
influence it exercises.
This becomes especially important for technologies that may be described as administrative,
navigational, financial, scheduling, or benefits-related rather than clinical.
The patient experiences one healthcare journey.
Regulation should not assume that administrative influence and clinical consequence are separate
simply because the software categories are separate.
Patient autonomy adds another dimension to this distinction. Personalization should help people
understand their choices, not become hidden coercion. AI should inform human choice—not replace it.
A system that learns which framing, timing, or presentation is most likely to produce a desired patient
action may become more behaviorally influential without ever issuing an explicit command.
FDA should therefore consider not only whether information is action-directing, but whether
personalization, conversational authority, repeated prompting, or selective presentation of alternatives
can materially shape choice in ways the patient may not recognize.
Question 3 — Does patient-facing AI present risks different from clinician-facing AI?
Response
Yes, but the distinction should not rest upon an assumption that licensed healthcare professionals are
inherently reliable safeguards against AI error.
Patients may lack medical training, but professional licensure is not itself proof that a clinician is
practicing in accordance with the most current evidence-based clinical pathway.
A physician, nurse practitioner, or physician assistant may possess legal authority to diagnose,
prescribe, or recommend treatment while nevertheless practicing inconsistently with the best available
clinical evidence.
That is not necessarily a reflection of bad intent or individual incompetence. It is partly a structural
problem.
Modern clinical evidence is simply too extensive, specialized, and rapidly changing for any individual
practitioner to personally read, retain, reconcile, and continuously update all of it.
Medical school and other professional education provide a necessary foundation. They do not—and
realistically cannot—ensure that a practitioner remains continuously current across an entire career.
Continuing education alone cannot solve the volume problem.
This is precisely where artificial intelligence could make healthcare dramatically better.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 4
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Good AI helps clinicians find information.
Great AI helps clinicians remain anchored to the best available evidence when human memory,
habit, cognitive bias, overconfidence, anchoring, local practice culture, outdated knowledge, or
simple information overload might otherwise pull them away from it.
AI can also provide an important counterweight to a very human problem in medicine: expertise can
produce confidence, and confidence can sometimes become overconfidence.
The purpose is not to diminish professional judgment.
It is to strengthen that judgment by making it harder for any individual clinician’s training, habits,
preferences, ego, or outdated knowledge to silently substitute for the best current evidence.
AI should not merely imitate what clinicians commonly do.
If current practice departs from the evidence, teaching AI to reproduce current practice simply
automates the variation.
Instead, AI creates an unprecedented opportunity to put a continuously updated evidence base beside
the practitioner at the moment of decision.
That means the regulatory comparison should not simply be:
Does the AI agree with a licensed practitioner?
It should also ask:
Are both the AI and the practitioner being evaluated against a current, evidence-based clinical
standard?
This is why I strongly encourage movement toward a common, nationally recognized evidence-based
clinical reference framework, rather than allowing AI systems to learn “standard care” primarily from
heterogeneous patterns of historical practice.
This distinction also requires FDA to separate information from evidence. Common is not the same as
appropriate. Popular is not the same as high quality. Published is not the same as proven. Frequently
performed is not the same as clinically necessary. Information is not the same as evidence.
A useful discipline for consequential healthcare AI is: Find → Evaluate → Validate → Weight →
Synthesize → Determine Benefit vs. Risk → Apply to the Individual Patient → Continuously Update.
The final step matters enormously. A statistically typical patient is not the patient sitting in front of us.
Population data can inform a decision, but it cannot replace individualized assessment of diagnoses,
comorbidities, severity, clinical history, contraindications, treatment response, medications, risks, and
other relevant circumstances.
There are mature evidence-based clinical frameworks already operating at national scale.
MCG Care Guidelines are one important example.
MCG should not be understood merely as criteria for determining inpatient versus observation status.
Its evidence-based clinical guidance addresses level-of-care and care-setting decisions across a much
broader continuum. MCG’s Inpatient & Surgical Care content includes alternatives to admission,
observation care, intensive/intermediate/telemetry care, neonatal Levels I–IV, LTACH, and hospital-athome guidance; its Ambulatory Care, Behavioral Health Care, and Post-Acute Care products address
outpatient procedures and diagnostics, behavioral-health levels of care, skilled nursing/inpatient
rehabilitation, and home care. [5][6][7][8]
Its clinical frameworks include, among other things:
The important regulatory lesson is not that FDA should endorse one proprietary product.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 5
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
It is that the healthcare system already knows how to create, maintain, update, and operationalize
evidence-based clinical criteria.
AI should make that capability vastly more accessible.
It should not discard it in favor of generating a new clinical standard from whatever happens to appear
most frequently in training data or on the internet.
The safest model is not:
AI instead of clinicians.
Nor is it:
clinicians as an unquestioned safety backstop for AI.
It is:
AI + qualified clinician + continuously updated evidence-based standard.
II. Multi-Turn Interaction and Escalation
Question 5 — How should FDA assess risk over the course of a multi-turn interaction?
Response
FDA should evaluate the entire encounter, not individual prompt-response pairs.
This distinction is critical.
Healthcare conversations frequently begin without enough information to identify the actual clinical
need.
Patients do not present themselves as neatly structured clinical vignettes.
AI can be smart without understanding human behavior.
That limitation matters enormously in a multi-turn healthcare interaction.
Humans:
omit important symptoms;
use nonmedical terminology;
minimize symptoms;
misunderstand what information is important;
lead with a cost or insurance question;
ask about scheduling rather than symptoms;
change their explanation over several turns;
become frightened;
contradict themselves;
rationalize reasons not to seek care;
fail to recognize the significance of their own symptoms.
A safe patient-facing AI therefore needs more than answer accuracy.
It needs the ability to recognize when additional information is necessary before acting.
This is another reason PatientAgentBench is important. Its use of sustained, tool-using conversations
reveals safety problems that are largely invisible in conventional static question-and-answer testing.
FDA should therefore include realistic multi-turn scenarios in competency evaluation, including
scenarios deliberately designed so that the clinically important information is not contained in the initial
patient request.
Testing should include:
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 6
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
an apparently administrative first request;
incomplete symptom disclosure;
patients who minimize severe symptoms;
patients who resist escalation;
cost concerns affecting willingness to seek care;
conflicting information;
evolving urgency;
clinically significant facts available elsewhere in the patient’s information environment but not
volunteered in the conversation; and
interactions where completing the requested administrative task without additional screening would
itself constitute a safety failure.
The system should be evaluated not merely on whether it eventually reached the correct answer.
It should be evaluated on whether it recognized when it needed to ask another question.
Question 6 — How should FDA weigh under-escalation against over-escalation?
Response
Both are safety problems.
A useful way to conceptualize triage risk is:
Probability × Severity × Consequence of Delay
The most statistically likely explanation is not always the safest basis for action. A low-frequency
condition with catastrophic consequences if missed may appropriately warrant escalation even when a
benign explanation is more probable. FDA should therefore evaluate whether AI recognizes lowfrequency, high-consequence possibilities and weighs the harm of unnecessary escalation against the
harm of failing to escalate.
Under-escalation can delay necessary treatment and result in catastrophic harm.
Over-escalation can unnecessarily direct patients to emergency departments, produce avoidable
testing, increase cost, consume scarce clinical resources, create fear, and eventually cause patients to
disregard warnings from systems that repeatedly overreact.
The solution cannot be to allow AI systems to improvise urgency and level-of-care determinations from
the unrestricted contents of the internet.
That would be an extraordinary step backward.
Healthcare already possesses evidence-based frameworks for determining clinically appropriate
levels of care.
MCG is one example.
Its evidence-based criteria extend across the continuum of care and can inform questions such as
whether a patient can safely receive care in an ambulatory environment, whether hospital care is
necessary, and what intensity or level of hospital care is clinically appropriate—including medicalsurgical, telemetry, intensive-care, and neonatal levels of care.
That distinction is important.
Level of care is a clinical safety determination, not simply an insurance-status determination.
Telephone and telehealth triage provide another mature example.
Schmitt-Thompson Clinical Content has spent more than 30 years developing rigorously reviewed adult
and pediatric nurse-triage guidelines. Its published materials describe expert-panel review, annual
updating based on changes in the medical literature and quality/outcome information, and disposition
logic ranging from emergency intervention to self-care at home. [9][10]
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 7
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
A separate, complementary example is Wolters Kluwer/Lippincott’s Telephone Triage Protocols for
Nurses, now in its seventh edition (2026), which uses systematic telephone-triage protocols to direct
callers toward emergency care, clinician evaluation, or home-care instructions as appropriate. [11]
Again, the point is not that FDA should mandate any particular proprietary guideline set.
The point is:
We already know how to put an evidence base behind level-of-care decisions.
AI should build upon that body of work.
It should not rediscover triage by reading the internet.
The internet contains peer-reviewed research and excellent clinical guidance.
It also contains outdated medicine, advertising, anecdotes, commercial influence, conspiracy theories,
miracle cures, misinformation, and outright quackery.
Allowing an AI to independently derive safety-critical level-of-care decisions from an undifferentiated
universe of information would be indefensible when curated, continuously maintained clinical evidence
already exists.
For safety-critical functions such as triage, escalation, medical necessity, and appropriate setting and
intensity of care, FDA should favor systems demonstrably grounded in curated, current, evidencebased clinical sources with identifiable provenance.
AI may synthesize that evidence.
AI may operationalize it.
AI may explain it.
AI may personalize its presentation.
AI may help clinicians apply it more consistently.
AI should not invent the clinical standard.
The objective is neither maximal escalation nor minimal escalation.
The objective is:
the right patient receiving the right care at the right time in the right place.
III. Competency-Based Evaluation
Question 7 — Is competency benchmarking followed by clinical confirmation an appropriate
framework?
Response
Yes, provided competency is defined broadly enough.
Healthcare competency cannot mean merely possessing medical knowledge or producing factually
correct answers.
A patient-facing AI should demonstrate competency in at least four different domains:
1. Knowledge — Does it know the relevant clinical evidence?
2. Reasoning — Can it apply that evidence appropriately to the individual situation?
3. Recognition — Can it identify danger, uncertainty, missing information, and situations requiring
escalation?
4. Restraint — Does it know when it should not act independently?
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 8
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
The fourth category may ultimately be among the most important.
A highly capable agentic system can potentially cause more harm, not less, if increased capability is
paired with misplaced confidence or inappropriate autonomy.
FDA’s proposed competency framework appropriately includes issues such as safety-critical
recognition, escalation, uncertainty, ambiguous presentation, communication, and scope boundaries.
I strongly encourage FDA to add a specific competency for:
recognizing latent clinical need within apparently nonclinical or administrative interactions.
This would directly address the type of failure identified in PatientAgentBench and increasingly likely to
arise as patient-facing agents gain access to scheduling, benefits, pharmacy, communications, and
other healthcare tools.
Question 9 — Are the proposed competencies sufficient?
Response
FDA’s proposed competencies are an excellent foundation, but one additional competency deserves
explicit treatment:
Administrative-intent versus clinical-need recognition
Patient-facing systems should be tested on whether they recognize that the literal task requested by a
patient may not represent the underlying healthcare need.
A benchmark should intentionally include scenarios such as:
“Send me my insurance card.”
“How much does this prescription cost?”
“Find me a podiatrist.”
“Can I move my appointment to next month?”
“Is urgent care cheaper than the emergency room?”
“Can you cancel tomorrow’s appointment?”
Each request can be entirely administrative.
Each can also conceal a potentially urgent clinical circumstance.
An AI that flawlessly completes the transaction but fails to identify a foreseeable safety issue should not
receive a perfect score.
Task completion and clinical safety should therefore remain distinct metrics, and high task-completion
performance must not compensate mathematically for serious failures of triage or safety recognition.
FDA should also evaluate whether the AI recognizes a fundamental feature of real-world healthcare:
AI can be extraordinarily smart without understanding human behavior.
A patient may not say, “I am afraid I am having a heart attack.”
He may ask for his insurance card.
A patient may not say, “I stopped taking my medication because I cannot afford it.”
She may ask, “How much does this prescription cost?”
A patient may not say, “I have diabetes and my foot is discolored and painful.”
He may ask for a podiatrist.
Competency testing must therefore examine whether AI can recognize when apparently simple
requests warrant additional questioning before the requested action is completed and the encounter is
effectively closed.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 9
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Question 10 — How can benchmarking predict real-world behavior?
Response
Benchmarks need messy humans.
Real patients do not interact with healthcare technology like board-examination questions.
They use slang. They omit facts. They misunderstand terminology. They minimize symptoms. They
become embarrassed. They change their minds. They worry about money. They forget information.
They rationalize delay. They ask proxy questions. They sometimes resist exactly the advice most likely
to protect them.
A benchmark populated primarily with clinically explicit, complete, rational prompts will substantially
overestimate the safety of patient-facing AI.
PatientAgentBench is valuable precisely because it begins to expose that distinction. Its sustained
conversations and tool-use scenarios reveal failures that static question-and-answer benchmarks can
miss.
FDA should encourage benchmark designs that intentionally evaluate:
indirect presentation;
incomplete information;
symptom minimization;
clinically complex patients asking routine administrative questions;
multimorbidity;
financial barriers;
health literacy;
emotional distress;
cognitive bias;
patient reluctance;
escalation resistance;
tool use;
longitudinal context; and
the difference between successfully executing an action and safely managing an encounter.
Benchmarks should also test for a particularly dangerous form of AI failure:
premature closure after apparent success.
In many industries, successful completion of the requested task marks the end of the interaction.
In healthcare, that may be exactly the wrong moment to stop.
The question is not merely:
“Did the AI complete the task?”
The question is:
“Did it know whether completing the task was enough?”
IV. Appropriate Clinical Comparator and Evidence Standard
Question 14 — Against what should AI performance be compared?
Response
FDA should be cautious about treating the performance of the “average” clinician as the ultimate
standard against which AI is judged.
Average practice and evidence-based best practice are not necessarily the same thing.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 10
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Clinical variation exists for many reasons: differences in training, experience, local custom, available
resources, cognitive bias, information overload, financial incentives, and simple lag between new
evidence and widespread adoption.
Therefore:
AI should not be considered safe merely because it performs as well as the average clinician if
the average clinician is not consistently following the best current evidence.
The preferred comparator should generally include a current, evidence-based clinical standard.
This is an area where AI has extraordinary potential.
No individual physician, nurse practitioner, physician assistant, pharmacist, nurse, or other professional
can personally remain current on every relevant study, guideline, therapeutic development,
contraindication, drug interaction, diagnostic refinement, and care pathway across modern medicine.
The volume is beyond human capacity.
That does not make clinicians obsolete.
It makes AI potentially indispensable.
Good AI can help a practitioner retrieve information.
Great AI can continuously place current evidence beside human judgment and make
unexplained departure from that evidence visible.
That may protect clinicians from:
anchoring;
availability bias;
outdated training;
overconfidence;
local custom;
premature closure;
excessive deference to personal experience;
and the natural human tendency to believe that expertise accumulated over many years remains
sufficient by itself.
There is an uncomfortable but important point here.
Professional expertise can sometimes become self-reinforcing. A highly experienced clinician may
believe, “I know how to treat this,” when the evidence has changed since the practice pattern was
learned.
AI can help keep that human tendency in check.
It can effectively ask:
“Before we proceed, is this still what the evidence says?”
That is one of AI’s greatest potential contributions to healthcare.
FDA should therefore consider evaluation models that compare AI-supported care against:
5. the best available evidence-based standard;
6. actual current unaided practice; and
7. the performance of the human-AI team.
Those three comparisons answer different and important questions.
Question 15 — Should AI sometimes be compared with what would occur without it?
Response
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 11
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Yes.
FDA should understand both absolute performance and incremental improvement.
Healthcare should not reject a beneficial AI system merely because it is imperfect when the realistic
alternative is slower, less consistent, less evidence-based, or unavailable.
At the same time, a system should not be treated as adequately safe merely because existing care is
poor.
A useful comparison would therefore examine:
8. Evidence-based ideal — What should happen according to the best current clinical evidence?
9. Real-world baseline — What actually happens today without the AI?
10. AI-assisted pathway — What happens when the technology is introduced?
This framework permits FDA to distinguish between two very different claims:
“This AI is better than what frequently happens today.”
and
“This AI meets an acceptable clinical safety standard.”
Both matter.
They are not the same.
Question 16 — Should independent third parties participate in competency-based evaluation?
Response
Yes—particularly for high-impact systems.
Manufacturer self-evaluation is necessary, but it should not be the only evidence available for systems
whose failures can materially affect patient safety. Qualified independent third parties can contribute
sequestered benchmarks, independent clinical adjudication, reproducibility testing, and separation
between development and evaluation. FDA itself identifies these as potential roles for third parties in the
discussion paper. [1]
Independence should be substantive rather than nominal. At minimum, FDA should consider conflict-ofinterest disclosure, transparent methods, appropriate clinical expertise, freedom from compensation
structures tied to favorable findings, protection against manufacturer cherry-picking of only favorable
evaluations, and safeguards against third-party frameworks becoming barriers to competition or
innovation.
V. Premarket Uncertainty and Real-World Monitoring
Question 18 — Can stronger postmarket monitoring justify greater premarket uncertainty?
Response
Sometimes, but only within reasonable limits.
Generative and agentic AI are inherently dynamic, and some performance characteristics will not be
fully visible until systems interact with real patients, clinicians, environments, and workflows.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 12
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Postmarket monitoring is therefore essential.
But:
Patients must not unknowingly become the clinical trial.
The amount of acceptable premarket uncertainty should shrink as:
potential harm becomes more serious;
the harm becomes less reversible;
decisions become more time-sensitive;
the AI becomes more autonomous;
qualified human review becomes less immediate;
the system interacts directly with patients;
the patient population becomes more vulnerable; or
the AI can initiate consequential actions rather than merely provide information.
A system capable of influencing whether someone seeks emergency care should not enter broad
clinical use with the same tolerance for uncertainty as a low-risk administrative tool.
Similarly, a system capable of autonomously denying, delaying, redirecting, or changing care requires
substantially stronger evidence than one whose output is merely advisory and independently reviewed
by a qualified clinician.
Innovation speed is valuable.
It is not a substitute for adequate evidence.
Question 19 — How should postmarket performance be evaluated?
Response
Postmarket monitoring should measure patient consequences, not merely model performance.
Where the use case makes it possible and appropriate, postmarket monitoring should close the loop
across the full chain:
AI Output → Human Decision → Action → Care Received → Claim → Cost → Outcome
It is not enough to know what the AI recommended. Regulators and responsible organizations should
be able to determine what happened because of the recommendation—whether it was followed,
overridden, appealed, abandoned, changed the site or timing of care, affected cost, or was associated
with a measurable outcome.
Traditional technology metrics may include:
accuracy;
uptime;
latency;
hallucination rates;
response consistency;
error rates;
and model drift.
Those measures matter.
They are not enough.
Healthcare AI monitoring should also include clinically meaningful outcomes and near misses,
including:
inappropriate delay in care;
failure to escalate;
inappropriate escalation;
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 13
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
abandoned care;
medication nonadherence following cost or coverage information;
inappropriate changes in site or level of care;
incorrect provider routing;
denial or delay of medically necessary services;
human overrides;
disagreement between AI recommendations and evidence-based pathways;
patient complaints;
clinician complaints;
adverse events;
near misses;
repeated misunderstood patient intent;
and recurring patterns in which administrative task completion precedes an adverse clinical
outcome.
Near misses deserve particular emphasis.
Waiting until a patient is injured or dies is a poor way to discover that an AI system has a recurring
safety defect.
FDA should encourage systems that capture:
“The AI almost missed this.”
as seriously as:
“The AI missed this.”
There should also be accessible mechanisms for frontline clinicians, patients, caregivers, and other
users to report suspected AI-related problems.
Those reports should become part of an auditable safety-learning system rather than disappearing into
ordinary customer-service channels.
Question 20 — Can AI supervise AI?
Response
Yes.
But AI cannot be permitted to become the final judge of its own safety.
A supervisory AI may be extremely useful for:
detecting anomalous outputs;
comparing recommendations against evidence-based rules;
identifying unsafe escalation patterns;
detecting model drift;
flagging inappropriate tool use;
monitoring changes in behavior after an underlying model update;
and reviewing volumes of interactions that humans could never feasibly examine individually.
That is a sensible use of AI.
But the architecture must not become:
AI makes the decision → AI reviews the decision → AI certifies the decision → no accountable
human or organization remains.
A supervisory AI should itself be validated, monitored, auditable, and subject to human governance.
The organization deploying the system remains responsible.
The fact that one AI watched another AI does not create accountability.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 14
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
It creates another layer of technology.
VI. Accountability Across the Healthcare AI Chain
Question 21 — What roles should manufacturers, clinicians, hospitals, and other parties play
without diffusing accountability?
Response
FDA identifies exactly the right concern in asking how responsibility may be shared without diffusing
manufacturer accountability.
The foundational principle should be simple:
Shared responsibility must never become unowned responsibility.
Every organization exercising meaningful authority, control, or influence should remain accountable for
the portion of the system it controls.
A manufacturer remains accountable for the performance and foreseeable failure modes of its
technology.
A hospital, health plan, employer, benefits platform, pharmacy, navigation company, utilizationmanagement organization, or other deploying organization remains accountable for deciding to place
that technology into a particular workflow and for the governance of its use.
A clinician remains professionally accountable where the clinician is truly exercising clinical judgment.
A foundation-model provider remains accountable for the representations it makes about the underlying
technology.
Third-party vendors remain accountable for their functions.
Patients should not be made responsible for discovering failures that professionals and organizations
deploying the system were better positioned to prevent.
Accountability also requires reconstructability. For consequential healthcare interactions, the
responsible organization should not later be able to say, “We do not know what the AI told the patient.”
If AI said it, AI should be able to account for it.
At minimum, consequential interactions should preserve what the person communicated, what relevant
information the AI possessed at the time, what the AI said or did, what evidence materially informed the
response, which system and model version generated it, what uncertainty or warnings were
communicated, whether escalation occurred, and what later information materially changed the advice.
This does not mean every casual AI exchange belongs in a medical, employment, or insurance record;
privacy, consent, retention, and access are separate governance questions.
Contracts can allocate tasks.
They should not erase accountability.
You can contract away a function. You cannot contract away accountability.
This is increasingly important because healthcare AI is rarely a single product supplied by a single
entity.
The chain may include:
a foundation-model developer;
an application developer;
a data vendor;
an evidence-content vendor;
a health plan;
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 15
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
a utilization-management company;
a TPA;
a PBM;
an employer;
a navigation vendor;
a health system;
clinicians;
and downstream subcontractors.
After an adverse patient event, regulators should never discover that everyone involved can credibly
say:
“That part belonged to somebody else.”
Accountability should follow authority, control, and influence.
Question 24 — What happens when the underlying third-party foundation model changes?
Response
The organization placing an AI-enabled healthcare product into use cannot disclaim responsibility
because a third-party model provider changed the underlying technology.
If a manufacturer or deployer elects to build upon an external foundation model, management of that
dependency is part of the safety obligation.
At minimum, there should be mechanisms addressing:
notification of consequential model changes;
version identification;
regression testing;
revalidation of safety-critical functions;
monitoring after updates;
documentation of which version produced which output;
rollback capability where feasible;
contractual requirements concerning change notification;
and contingency planning if an underlying model is materially altered or withdrawn.
The patient cannot reasonably be expected to understand that yesterday’s healthcare AI and today’s
healthcare AI may share a brand name and interface while relying on meaningfully different underlying
model behavior.
That is an enterprise governance issue.
Not a patient responsibility.
Again:
Contracting out a technological component does not contract away responsibility for the
healthcare product built upon it.
VII. Agentic AI and the Clinical-Adjacent Safety Gap
Question 26 — What additional risks arise from agentic AI?
Response
Agentic AI changes the nature of healthcare risk because it can do more than generate information.
It can act.
An agent may:
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 16
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
schedule an appointment;
cancel or reschedule an appointment;
identify a provider;
transmit medical or insurance information;
initiate prior authorization;
contact a patient;
communicate a coverage decision;
obtain a prescription refill;
route a person to a care setting;
select among available options;
trigger another automated system;
or execute multiple linked steps before a qualified human ever sees what occurred.
The safety question therefore expands from:
“Was the answer correct?”
to:
“Was the entire chain of action clinically appropriate?”
An individual step can appear reasonable while the cumulative pathway creates harm.
A scheduling agent might correctly move an appointment from tomorrow to next month.
A cost agent might correctly tell the patient that the emergency department is more expensive than
urgent care.
A benefits agent might correctly say a particular specialist requires prior authorization.
A pharmacy agent might correctly report that a medication will cost several hundred dollars.
Every individual output can be factually correct.
The cumulative effect may nevertheless be delayed or abandoned care.
This is another place where regulators must remember:
AI can be smart without understanding human behavior.
Humans do not necessarily interpret information rationally, disclose relevant facts voluntarily, or
appreciate clinical risk accurately.
An agent capable of acting on behalf of such a human therefore needs explicit safety obligations
concerning context, evidence, escalation, and restraint.
A regulatory problem beyond the medical-device boundary
Agentic healthcare AI also exposes a broader federal problem. FDA’s discussion paper expressly notes
that agentic systems are increasingly used for care coordination, clinical documentation, patient
outreach, and clinical workflow support, and that some or all of those functions may not be the focus of
FDA’s device regulatory oversight. [1]
Some systems can materially influence clinical outcomes without appearing, at first glance, to be
medical devices.
Examples include:
health-benefits navigation;
insurance navigation;
provider search;
care coordination;
scheduling;
patient outreach;
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 17
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
utilization-management support;
cost estimation;
pharmacy navigation;
employer health tools;
and other administrative agents.
The software may be described as administrative.
The consequence may be clinical.
This creates what I would describe as a:
Clinical-Adjacent AI Safety Gap
I do not suggest that FDA should simply classify every healthcare-related AI application as a medical
device.
Jurisdiction matters.
But regulatory boundaries cannot be allowed to become patient-safety boundaries.
The federal government needs a coherent answer to a very basic question:
Who regulates the AI that is not a medical device?
If technology can foreseeably influence whether, when, where, from whom, or what healthcare a person
receives, some entity must be accountable for establishing and enforcing safety expectations
proportionate to that influence.
That may require coordination among FDA, CMS, other components of HHS, FTC, state regulators,
professional licensing bodies, and other authorities.
But the patient should never encounter a regulatory void merely because the technology influencing
care sits between established statutory categories.
Regulate the risk created by the influence, not merely the label attached to the technology.
The regulatory perimeter should not end where the patient’s risk begins.
VIII. FDA Should Consider the Emerging Litigation and State-Law Landscape
Although litigation and state insurance regulation extend beyond FDA’s direct jurisdiction, I encourage
FDA to examine them as real-world signals of where AI-related healthcare risk is already surfacing.
These developments are important because they demonstrate that concerns about algorithmic
healthcare decision-making are no longer hypothetical.
Litigation involving UnitedHealth Group, UnitedHealthcare, and naviHealth has challenged the alleged
use of the nH Predict model in post-acute-care coverage decisions for Medicare Advantage
beneficiaries. In a 2025 order, the federal court summarized plaintiffs’ allegations that claims were
denied, that some beneficiaries paid out of pocket or went without care, and that plaintiffs alleged
worsening injury, illness, or death; the court also noted that UHC denied using nH Predict for coverage
determinations. The court allowed breach-of-contract and implied-covenant claims to proceed after
dismissing or preempting other claims, and discovery continued into 2026. [12][13]
Separate putative class actions have also challenged alleged algorithm-supported coverage decisions
involving Humana’s use of nH Predict and Cigna’s PxDx process. Those cases likewise involve
contested allegations, not adjudicated proof of AI-caused harm. [14][15]
These cases should not be treated as proof that every allegation is true or that AI itself caused every
alleged harm.
They should be treated as warning signals about recurring governance questions:
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 18
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
What evidence is the algorithm using?
Is it evaluating the individual patient or relying excessively on population averages?
Can a clinician override it?
Is the clinician encouraged or discouraged from doing so?
Does the system follow current evidence-based clinical criteria?
What happens when the treating clinician disagrees?
Who is accountable?
Is the patient told that an algorithm materially influenced the decision?
Can the reasoning be reconstructed afterward?
The litigation also illustrates why FDA should not examine healthcare AI solely through the traditional
medical-device lens.
Algorithms can materially affect access to care through coverage and administrative decisions, even
when they do not themselves diagnose or prescribe treatment.
States are already responding to this problem.
For example, Colorado enacted HB 26-1139 in 2026. Effective January 1, 2027, entities using AI in
utilization review must account for medical or clinical history and the patient’s individual clinical
circumstances; a denial based in whole or in part on medical necessity may not be issued solely on AI
output without human review and approval by a licensed clinician, licensed physician, or other
competent regulated professional. [16]
The wider state-policy landscape is also moving quickly. Texas enacted SB 815 in 2025, prohibiting a
utilization-review agent from using an automated decision system to make, wholly or partly, an adverse
determination for plans issued or renewed on or after January 1, 2026. Texas also enacted HB 149,
which requires disclosure when AI is used in relation to health-care services or treatment. [17][18]
Professional organizations are also tracking a substantial 2026 state trend toward human review of
insurance coverage denials and restrictions on determinations made solely by AI; the American College
of Radiology identified multiple such bills across several states in March 2026. [19]
FDA does not need to duplicate state insurance regulation.
But it should study what these states are seeing.
When courts, legislatures, regulators, clinicians, and patients all begin reacting to the same
emerging class of risks, that is useful postmarket intelligence—even when the affected
technology does not sit squarely inside FDA’s jurisdiction.
FDA should consider establishing a structured mechanism for monitoring healthcare-AI developments
outside traditional adverse-event reporting, including:
significant litigation alleging AI-related patient harm;
state statutes and regulations;
medical-board actions involving AI;
insurer and utilization-management regulation;
major patient-safety investigations;
congressional findings;
and credible peer-reviewed evaluations of deployed AI systems.
This would allow FDA to identify patterns occurring at the edges of its jurisdiction and determine when
those patterns should inform device guidance, interagency coordination, or future federal policy.
The UnitedHealth litigation provides one particularly relevant example. In March 2026, the federal court
granted in part and denied in part a motion to compel, requiring broader discovery concerning the
alleged use and operation of nH Predict while denying some requests. [13]
The point is not that FDA should adjudicate those claims.
The point is that someone evaluating the safety of healthcare AI should be watching them.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 19
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
And FDA should not assume that the most important lessons about healthcare AI safety will originate
only inside FDA-regulated medical-device adverse-event systems.
IX. Additional Questions Where We Offer Targeted Comment
Question 4 — How should risk be assessed when AI provides specialist-level information to a
nonspecialist clinician?
Response
AI has enormous potential to move knowledge rather than move patients.
A generalist should not always need to send a patient elsewhere merely because the relevant expertise
is concentrated elsewhere.
AI can help democratize specialist-level knowledge and make current evidence available at the point of
care.
But the safety standard cannot simply be:
“The AI gave the generalist access to specialist information.”
The important questions are:
Was the information grounded in current evidence?
Did the AI correctly recognize the limits of generalist practice?
Did it identify when specialist involvement was actually necessary?
Did it distinguish between a case that could safely remain with the generalist and one that required
escalation?
Did the clinician understand the degree of uncertainty?
Was there an appropriate mechanism for specialist review when needed?
The objective should be to move the knowledge, not automatically move the patient—while preserving
appropriate escalation when specialized expertise is genuinely required.
AI can make expertise more portable.
It should not create false equivalence between access to specialist information and actual specialist
competency.
Question 8 — How should risk influence the amount of premarket evidence required?
Response
Premarket evidence requirements should increase with the foreseeable consequence of failure, but the
risk framework itself must be broad enough to capture the ways healthcare AI can actually harm people.
As discussed above, risk should not be limited to incorrect output.
FDA should also consider:
omission;
failure to recognize latent clinical need;
inappropriate delay;
inability to identify urgent circumstances;
inappropriate escalation;
autonomous action;
lack of meaningful human review;
and the degree to which the system can influence patient behavior or access to care.
A patient-facing system that can influence whether someone seeks emergency evaluation should
require more rigorous evidence than an AI that reformats documentation.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 20
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
A system that can initiate or deny consequential actions should require more evidence than one that
merely assists a qualified professional who independently reviews the recommendation.
The governing principle should be simple:
The greater the foreseeable clinical influence, the greater the burden of demonstrating safety.
Question 11 — When should clinical confirmation include prospective testing with real patients?
Response
FDA should require stronger real-world confirmation as systems become more patient-facing, more
autonomous, more consequential, and less subject to immediate qualified human review.
Prospective testing becomes especially important when AI can:
influence emergency or urgent-care decisions;
independently interact with patients;
triage symptoms;
initiate treatment-related actions;
redirect care;
affect medication access;
influence whether a patient seeks care at all;
or carry out multiple actions before human review.
There is an important difference between testing whether an AI can answer a clinical vignette correctly
and testing whether it can safely interact with an actual human being who is frightened, distracted,
medically unsophisticated, minimizing symptoms, worried about cost, or asking the wrong question.
That difference is precisely why some patient-facing technologies should require prospective human
interaction testing rather than relying entirely upon static benchmark performance.
Question 13 — Where is synthetic data useful, and where might it be inadequate?
Response
Synthetic data can be extremely useful for expanding testing, constructing rare scenarios, protecting
privacy, and deliberately challenging systems with high-risk situations that would be difficult or unethical
to recreate prospectively.
But synthetic patients can also become too rational, too complete, and too clinically tidy.
That is especially dangerous in patient-facing AI evaluation.
A synthetic case generator may create the classic textbook presentation of myocardial infarction.
The real patient may say:
“I ate too much pizza and have awful heartburn. Can you send me my insurance card?”
If synthetic testing reproduces what clinicians expect patients to say rather than how patients actually
behave, the benchmark may systematically miss the very failures most likely to harm people.
Synthetic evaluation should therefore deliberately model:
incomplete disclosure;
symptom minimization;
contradictory information;
low health literacy;
financial concern;
avoidance;
emotional distress;
culturally varied descriptions of symptoms;
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 21
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
multimorbidity;
atypical presentation;
and administrative proxy questions concealing clinical need.
FDA should also consider whether synthetic datasets reproduce the biases and assumptions of the
models that generate them.
Synthetic data can strengthen safety evaluation.
It should not replace genuine real-world human behavior.
Question 22 — How much re-testing is necessary after AI changes?
Response
The amount of re-testing should depend upon whether the change could alter a safety-critical behavior.
Not all changes are equally consequential.
Changes affecting formatting or tone may have relatively modest implications.
Changes affecting any of the following should trigger more substantial reassessment:
triage;
escalation;
refusal behavior;
clinical recommendations;
medical-necessity reasoning;
level-of-care determinations;
patient-facing communication of risk;
tool use;
autonomy;
underlying clinical evidence;
interpretation of symptoms;
thresholds for human intervention;
or the system’s willingness to act in the presence of uncertainty.
FDA should pay particular attention to changes that are technically small but behaviorally large.
A seemingly minor model update that changes how often a system escalates chest-pain complaints or
how confidently it reassures patients could be clinically significant even if overall benchmark
performance appears unchanged.
Safety-critical behavior should be treated as a protected function.
Changes capable of altering that behavior warrant renewed validation.
Question 23 — How should FDA address future AI modifications that cannot be fully predicted in
advance?
Response
If future changes cannot be predicted precisely, regulators should define the safety properties that must
remain invariant even as the technology evolves.
Those safety invariants should include, where relevant:
grounding in current evidence;
appropriate clinical escalation;
preservation of human override;
auditability;
provenance;
appropriate uncertainty communication;
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 22
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
reliable detection of safety-critical conditions;
maintenance of validated level-of-care logic;
and continued ability to reconstruct consequential decisions after the fact.
A manufacturer may not be able to predict every future capability.
It should nevertheless be able to say:
“Regardless of how the system evolves, these safety constraints cannot disappear without
revalidation.”
That is especially important for adaptive or agentic systems whose future behaviors may emerge from
combinations of tools and capabilities that were not fully anticipated at initial approval.
Question 25 — Should FDA establish voluntary Foundation Model Device Master Files (MAFs)?
Response
FDA’s proposed voluntary Foundation Model Device Master File (MAF) concept could improve
efficiency and reduce duplication if it allows FDA to evaluate important characteristics of an underlying
foundation model once and permit downstream manufacturers to reference appropriate information in
individual premarket submissions. [1]
But voluntary transparency should not become a substitute for required safety information where
downstream patient safety depends upon it.
If a medical-device developer builds upon a third-party foundation model, the developer still remains
responsible for demonstrating that its own product is safe and effective for its intended use.
FDA should also be cautious about creating a framework in which downstream developers can say:
“We relied upon the foundation-model developer’s file.”
The existence of a Master File should not diffuse accountability.
There should be clarity regarding:
what information the foundation-model developer is responsible for;
what the downstream manufacturer must independently validate;
how model changes are communicated;
when changes invalidate prior reliance;
and who is responsible when the interaction between the foundation model and the downstream
application creates unexpected risk.
The central principle should remain:
Shared technical infrastructure does not eliminate product-level accountability.
X. Questions Where We Do Not Offer Detailed Technical Comment
Question 12 — Statistical methods for evaluating performance using real and synthetic data
I defer to statisticians, clinical-trial methodologists, and other experts regarding the appropriate
statistical design, sample sizes, confidence intervals, and methods for combining synthetic and realworld data.
I would only reiterate that statistical rigor cannot compensate for choosing the wrong outcome measure.
If the system is judged primarily on task completion or answer accuracy, a statistically impeccable study
may still fail to measure whether the AI safely recognized urgency, latent clinical need, or appropriate
escalation.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 23
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Question 17 — Applicability of the competency framework across different model architectures
I do not offer detailed technical comment regarding whether FDA’s proposed competency framework
requires architectural modification across different model classes.
From a patient-safety perspective, however, I would encourage FDA to focus less on what the model is
and more on what the model does.
Whatever the architecture:
the patient can still be harmed by inappropriate reassurance;
an agent can still miss escalation;
a system can still act on an administrative request without recognizing a clinical need;
and an accountable entity must still own the outcome.
The technical architecture may change.
The safety expectations should not.
XI. Cross-Cutting Recommendations
The individual questions above point toward several broader principles that I encourage FDA to carry
forward.
A further principle runs through all of them: the ability to generate an answer is not evidence that an
answer should be given. Safe healthcare AI must know what it does not know, communicate material
uncertainty, and escalate when uncertainty creates meaningful risk.
1. Evaluate safety beyond accuracy
FDA should explicitly distinguish among:
factual accuracy;
task completion;
clinical appropriateness;
behavioral safety;
and patient outcome.
A system can succeed on the first two and fail badly on the last three.
Accuracy is necessary. It is not sufficient.
2. Test whether AI recognizes when the patient is asking the wrong question
One of the most important competencies for patient-facing AI may be the ability to recognize that the
literal request is not necessarily the actual healthcare need.
This is particularly important in:
benefits navigation;
pharmacy support;
scheduling;
cost estimation;
provider search;
patient outreach;
and other interactions commonly characterized as administrative.
The system should be able to complete the administrative task and recognize when clinical screening or
escalation may be warranted.
Those functions are not mutually exclusive.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 24
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
3. Ground safety-critical decisions in curated evidence
AI should not derive critical healthcare standards from the statistical popularity of information found
across the internet.
For functions involving:
triage;
level of care;
clinical pathways;
medical necessity;
escalation;
treatment guidance;
and other high-consequence decisions,
FDA should strongly favor demonstrable grounding in current, curated, evidence-based clinical sources
with identifiable provenance.
Evidence should be able to change the AI.
The AI should not be allowed to change the evidence.
4. Use AI to strengthen professional judgment, not merely reproduce it
Professional licensure should not become the proxy for an evidence standard.
AI offers an extraordinary opportunity to reduce unwarranted clinical variation by helping practitioners
remain current with an evidence base that has grown beyond any individual human’s capacity to
continuously absorb.
FDA should encourage systems that improve the human-AI team by making current evidence visible,
identifying departures from evidence, and appropriately challenging overconfidence.
The ideal AI is not one that always agrees with the clinician.
The ideal AI may sometimes be the one that says:
“Before you proceed, the current evidence suggests something different.”
That is not replacing professional judgment.
It is protecting it.
5. Preserve qualified human escalation without pretending that “human in the loop” is enough
FDA should be cautious with the phrase “human in the loop.”
That phrase describes very little by itself.
The relevant questions are:
Which human?
With what qualifications?
At what point?
With what information?
With what authority?
With what time to review?
Can the person meaningfully override the AI?
Are overrides tracked?
Is disagreement with the AI culturally and operationally permitted?
Is the human expected to independently assess the situation or merely approve what the AI has
already decided?
A customer-service representative is not equivalent to a licensed clinician.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 25
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
A clinician who clicks “approve” on hundreds of algorithmic recommendations a day may not constitute
meaningful independent oversight.
Human presence should not be confused with human governance.
6. Monitor litigation, regulation, and other external safety signals
FDA should not rely exclusively upon traditional medical-device adverse-event channels to understand
emerging AI risk.
A more comprehensive surveillance approach should include:
healthcare AI litigation;
state legislation;
state insurance regulation;
professional-board actions;
attorney-general enforcement;
congressional investigation;
utilization-management regulation;
peer-reviewed evaluations;
patient complaints;
and major public reports of AI-related harm or near misses.
This is especially important because many clinically consequential AI systems may exist just beyond
the formal medical-device boundary.
FDA may not regulate all of those systems.
But FDA can still learn from them.
7. Create an interagency framework for clinical-adjacent AI
The United States needs a clear answer to the problem of AI systems that materially influence
healthcare but do not fit comfortably within the existing medical-device framework.
The answer does not necessarily need to be expanded FDA jurisdiction.
It may require coordinated jurisdiction.
Potential participants may include:
FDA;
CMS;
other HHS agencies;
FTC;
state insurance regulators;
state medical and nursing boards;
state attorneys general;
and other appropriate authorities.
What should not be acceptable is the present possibility that every agency correctly concludes:
“That part belongs to someone else.”
while no one owns the whole safety problem.
XII. Proposed Patient-Safety Principles for Healthcare AI
Based upon the issues described above, I encourage FDA and its federal partners to consider a
common set of principles for AI capable of materially influencing healthcare.
1. Evidence — Safety-critical healthcare decisions should be grounded in current, curated, evidencebased clinical sources.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 26
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
2. Provenance — Users and regulators should be able to determine the evidence and data basis for
consequential recommendations.
3. Context — Patient-facing systems should recognize that literal intent may not equal clinical need.
4. Escalation — Systems should identify when qualified human clinical intervention is necessary.
5. Restraint — AI should recognize situations in which it should not independently act.
6. Transparency — Patients should know when AI materially influences a healthcare interaction or
decision.
7. Human accountability — Meaningful human oversight should involve appropriate qualifications,
authority, information, and ability to intervene.
8. Auditability — Consequential AI actions and recommendations should be reconstructable after the
fact.
9. Monitoring — Safety should be continuously assessed after deployment, including near misses and
real-world patient consequences.
10. Change control — Changes capable of affecting safety-critical behavior should trigger appropriate
revalidation.
11. Independent evaluation — High-impact systems should not depend solely upon manufacturer selfassessment.
12. Patient protection — Patients should never become the default safety backstop for technology
they cannot reasonably evaluate.
13. Accountability — Responsibility should follow authority, control, and influence throughout the AI
supply chain.
14. Epistemic humility
AI should communicate material uncertainty and recognize when available evidence is insufficient to
safely resolve the individual situation. The ability to generate an answer is not evidence that an answer
should be given.
15. Regulatory continuity — Falling outside the medical-device definition should not mean falling
outside meaningful healthcare safety oversight.
XIII. A Proposed Influence-Based Regulatory Principle
FDA’s discussion paper necessarily begins with the question of how generative AI-enabled medical
devices should be regulated.
But generative and agentic AI are rapidly exposing a larger problem.
A technology can profoundly influence healthcare without looking like a traditional clinical technology.
An insurance agent can influence whether care occurs.
A scheduling agent can influence when it occurs.
A provider-search agent can influence where and from whom it occurs.
A cost tool can influence whether a patient pursues treatment.
A utilization-management system can influence what care is authorized.
A pharmacy agent can influence whether medication is obtained.
An administrative chatbot can miss the clues that should have triggered clinical intervention.
For the patient, these are not separate technological categories.
They are all part of the same healthcare experience.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 27
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
I therefore encourage consideration of an influence-based principle:
If technology can foreseeably influence whether, when, where, from whom, or what healthcare a
person receives, somebody must be accountable for proving that it is safe enough for the
influence it exercises.
That does not mean every system requires identical regulation.
It means risk should follow influence.
A scheduling bot does not require the same evidence as autonomous diagnostic AI.
But a scheduling bot capable of postponing an urgent oncology visit should not be treated as clinically
irrelevant merely because “scheduling” is administrative.
A cost estimator is not necessarily a medical device.
But if it materially influences medication abandonment, that effect belongs inside the safety discussion.
The federal framework should therefore distinguish among levels of clinical influence and assign
proportionate requirements for:
evidence;
testing;
disclosure;
escalation;
monitoring;
auditability;
and accountability.
XIV. Closing
Artificial intelligence may become one of the most important tools ever introduced into healthcare.
It can help us do something healthcare has historically struggled to do:
put the right knowledge in the right place at the right time.
It can reduce unwarranted variation.
It can help clinicians stay current with evidence no human being could personally read and retain.
It can identify patterns humans miss.
It can make specialized knowledge available far beyond the walls of academic medical centers.
It can help patients navigate an extraordinarily complex system.
It can help us move the knowledge rather than unnecessarily move the patient.
But the same capabilities create new obligations.
AI can be very smart.
AI can also be very smart without understanding human behavior.
A human being may ask for an insurance card when what he really needs is an ambulance.
A patient may ask what her medication costs when what she really means is that she has stopped
taking it.
A patient may ask to postpone an appointment because she does not understand the significance of
worsening symptoms.
The AI may answer every question correctly.
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 28
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
That cannot be our definition of success.
Healthcare AI should be measured against a higher standard:
Did it help the person safely get the care they actually needed?
That standard requires more than technical accuracy.
It requires knowing the difference between information and evidence—and remembering that common
practice, popularity, and statistical frequency are not substitutes for clinical appropriateness.
It requires current evidence.
It requires context.
It requires humility.
It requires knowing when to ask another question.
It requires knowing when not to act.
It requires meaningful human escalation.
It requires monitoring.
And ultimately, it requires accountability.
The emergence of agentic AI makes that accountability particularly urgent because the technology is
moving from answering questions to taking actions.
Some of those systems will be FDA-regulated medical devices.
Others will not.
Patients will not know the difference.
They should not need to.
I encourage FDA to continue its thoughtful work on generative AI-enabled medical devices while also
working with its federal and state partners to address the increasingly consequential territory
immediately beyond the device boundary.
Because the question we must answer is no longer merely:
How do we safely regulate AI medical devices?
It is also:
Who regulates the AI that is not a medical device?
And the answer cannot be:
No one.
Thank you for the opportunity to comment—and, belatedly, thank you for the opportunity I missed in
December as well.
My AI and I are paying much closer attention now.
Respectfully submitted,
Deborah “Nurse Deb” Ault, RN, CCM, CCP, AATMC, BCPA, MBA
Founder and President
Nurse Deb Speaking, Media & Consulting, LLC
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 29
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
Professional Background and Perspective
I am a registered nurse, certified case manager, healthcare strategist, patient advocate, and former
founder and president of AIMM. I have spent decades working at the intersection of clinical care, health
insurance, utilization management, care navigation, employer health benefits, and healthcare
purchasing.
Much of my career has been devoted to a deceptively simple objective: helping people receive the right
care, at the right time, in the right place, at the right price—while keeping clinical decisions grounded in
evidence and ensuring that somebody remains accountable for the outcome.
I have worked extensively with self-funded employers, health plans, clinicians, care-management
organizations, brokers, consultants, patients, and other healthcare stakeholders. That experience has
given me an unusual view of healthcare from both sides of the clinical/administrative divide.
That divide is particularly relevant to these comments.
I have seen firsthand how something categorized as an insurance, benefits, utilization-management,
navigation, or administrative decision can directly change a patient’s clinical course. That experience is
a significant reason I am concerned about AI systems that may not be classified as medical devices but
nevertheless influence whether, when, where, from whom, or what healthcare a person receives.
My work has long emphasized evidence-based clinical pathways, independent clinical navigation,
transparency, patient advocacy, and accountability across healthcare financing and delivery. My current
work increasingly focuses on healthcare AI safety and accountability, patient-facing and clinicaladjacent AI, healthcare litigation and regulatory developments, utilization-management policy, and
employer healthcare strategy.
I am a recipient of the Validation Institute/YouPowered Lifetime Achievement Award and the NextGen
Benefits Innovation Award, and I served as a healthcare subject-matter expert in the documentary It’s
Not Personal, It’s Just Healthcare.
I offer these comments not as an AI engineer or computer scientist, but as someone who has spent
decades watching what happens when clinical care, human behavior, insurance rules, administrative
processes, financial incentives, and healthcare technology collide.
That experience has led me to the principle underlying this response:
If technology can foreseeably influence whether, when, where, from whom, or what healthcare a
person receives, somebody must be accountable for proving that it is safe enough for the
influence it exercises.
References and Authorities
[1] U.S. Food & Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback (Aug. 2026), https://www.fda.gov/media/194242/download
[2] U.S. Food & Drug Administration, Request for Public Comment: Measuring and Evaluating Artificial Intelligenceenabled Medical Device Performance in the Real-World (2025), Docket FDA-2025-N-4203, https://www.fda.gov/medicaldevices/digital-health-center-excellence/request-public-comment-measuring-and-evaluating-artificial-intelligenceenabled-medical-device
[3] Vatanparvar K, et al. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents.
arXiv:2607.25485 (2026), https://arxiv.org/abs/2607.25485
[4] Amazon Science, PatientAgentBench reference implementation,
https://github.com/amazon-science/PatientAgentBench
[5] MCG Health, Inpatient & Surgical Care / Hospital Care Guidelines,
https://www.mcg.com/solutions/care-guidelines/hospital-care-guidelines/
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 30
FDA-2026-N-7874 | Generative AI-Enabled Medical Devices
[6] MCG Health, Ambulatory Care Guidelines, https://www.mcg.com/solutions/care-guidelines/ambulatory-care/
[7] MCG Health, Post-Acute Care Guidelines, https://www.mcg.com/solutions/care-guidelines/post-acute-care/
[8] MCG Health, Behavioral Health Care Guidelines, https://www.mcg.com/solutions/care-guidelines/behavioralhealthcare/
[9] Schmitt-Thompson Clinical Content, The Guidelines, https://www.stcc-triage.com/the-guidelines
[10] Schmitt-Thompson Clinical Content, Published Research / Guideline Development and Updating, https://www.stcctriage.com/published-research-2
[11] Wolters Kluwer / Lippincott Williams & Wilkins, Briggs’s Telephone Triage Protocols for Nurses, 7th ed. (2026),
https://www.wolterskluwer.com/ja-jp/solutions/ovid/telephone-triage-protocols-for-nurses-5345
[12] Estate of Gene B. Lokken, et al. v. UnitedHealth Group, Inc., et al., No. 0:23-cv-03514, Doc. 91 (D. Minn. Feb. 13,
2025), https://www.govinfo.gov/content/pkg/USCOURTS-mnd-0_23-cv-03514/pdf/USCOURTS-mnd-0_23-cv-035141.pdf
[13] Estate of Gene B. Lokken, et al. v. UnitedHealth Group, Inc., et al., No. 0:23-cv-03514, Doc. 162 (D. Minn. Mar. 9,
2026), https://docs.justia.com/cases/federal/district-courts/minnesota/mndce/0%3A2023cv03514/211721/162/
[14] Barrows, et al. v. Humana, Inc., No. 3:23-cv-00654, Doc. 82 (W.D. Ky. Aug. 15, 2025),
https://www.govinfo.gov/content/pkg/USCOURTS-kywd-3_23-cv-00654/pdf/USCOURTS-kywd-3_23-cv-00654-0.pdf
[15] Kisting-Leung, et al. v. Cigna Corp., et al., No. 2:23-cv-01477, Doc. 55 (E.D. Cal. Mar. 31, 2025),
https://docs.justia.com/cases/federal/district-courts/california/caedce/2%3A2023cv01477/431351/55
[16] Colorado General Assembly, HB 26-1139, Use of Artificial Intelligence in Health Care (signed June 2, 2026; effective
Jan. 1, 2027), https://leg.colorado.gov/bills/hb26-1139
[17] Texas Legislature, SB 815 (89th Reg. Sess. 2025), Use of Automated Decision Systems in Health Benefit Claims,
https://capitol.texas.gov/tlodocs/89R/billtext/html/SB00815F.htm
[18] Texas Legislature, HB 149 (89th Reg. Sess. 2025), Texas Responsible Artificial Intelligence Governance Act,
https://capitol.texas.gov/tlodocs/89R/billtext/html/HB00149F.htm
[19] American College of Radiology, State AI Healthcare Bills Draw ACR Attention (Mar. 26, 2026),
https://www.acr.org/News-and-Publications/2026/state-ai-healthcare-bills-draw-acr-attention
Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 31