American College of Allergy, Asthma & Immunology
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
See attached file(s)
Attachment
SUBMITTED ELECTRONICALLY VIA WWW.REGULATIONS.GOV
September 25, 2026
The Honorable Kyle Diamantas, J.D.
Acting Commissioner
Food and Drug Administration
Department of Health and Human Services
Attention: FDA-2026-N-7874
10903 New Hampshire Ave.
Silver Spring MD 20993
RE: ACAAI’s Comments on the FDA's Digital Health Center of Excellence (DHCoE)
recent discussion paper and request for feedback, Considerations for the
Regulation of Generative AI-Enabled Medical Devices
Dear Acting Commissioner Diamantas:
The American College of Allergy, Asthma and Immunology (“ACAAI”) and its Advocacy Council
appreciate the opportunity to submit comments on the FDA's Digital Health Center of Excellence
(DHCoE) recent discussion paper and request for feedback, Considerations for the Regulation
of Generative AI-Enabled Medical Devices. The ACAAI represents the interests of more than
6,500 allergists-immunologists and allied health professionals. Our members provide patient
services across a variety of settings ranging from small or solo physician offices to large
academic medical centers. As a result, our membership brings a broad perspective on how AIenabled medical devices can affect the quality and safety of care provided to our patients, and
the importance of regulating these devices.
Below are our complete responses to each question raised in your discussion paper.
Question 1. Does the two-axis risk framework appropriately capture the dimensions most
relevant to risk? What additional dimensions should be represented?
We support the two-axis framework as a good starting point, but we want to see a few additions.
First, reversibility of a resulting action should be its own explicit dimension, not folded into
consequences. In allergy and immunology, a delayed or inaccurate recommendation about
anaphylaxis recognition is far less reversible than a delayed recommendation about
environmental allergen avoidance counseling, even if both might otherwise be scored at a similar
severity level.
Second, whatever action a device takes or recommends, clinicians need a reliable way to verify
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200
AdvocacyCouncil@acaai.org
what the device actually did, or is recommending, and to reverse or override that action if needed.
This matters most for agentic or higher autonomy functions, but it also applies to informational
tools that influence downstream decisions.
Third, we ask CDRH to weight the time pressure of the deployment setting. A tool supporting real
time triage of a possible anaphylactic reaction operates under very different constraints than one
supporting long term asthma control planning, and premarket evidence expectations should
reflect that. Finally, traceability to primary source material, such as current practice parameters,
should be a core requirement given how quickly our field's guidance evolves, for example with
new biologics or food OIT protocols.
Question 2. How to account for the spectrum between 'non-directive' and 'action-directing'
informational functions, and what characteristics should modify risk?
We agree that directiveness sits on a continuum and should be judged by substance rather than
by surface language like recommend or should. In our specialty, a tool that turns a symptom
pattern into a specific action, for example telling someone to start an antihistamine at a certain
dose versus simply noting that antihistamines are commonly used for these symptoms, should be
treated as action directing regardless of any hedging language or disclaimers attached to it. We
ask CDRH to require manufacturers to document, as part of intended use, where a given output
type is meant to sit on this continuum, and to test specifically for unintended drift toward
directiveness. This is especially relevant for allergy specific tools that field open ended questions
about medication dosing, epinephrine use, or immunotherapy.
Question 3. When might patient-facing delivery of clinical information present different or higher
risk than HCP-facing delivery, and what safeguards might mitigate this while preserving patient
empowerment?
Patient facing GenAI tools carry real added risk in our field, and this deserves more weight than
the current framework implies. Physicians are trained specifically to draw out a complete and
nuanced history because patients often do not spontaneously report every relevant detail, and
they do not always connect a symptom to their presenting concern(s). A child with asthma might
complain of a sore throat when what they are really describing is trouble breathing. An adult with
asthma might say their allergies are acting up but to them it means their asthma is flaring. A GenAI
tool talking directly with a patient does not have that trained ear and will be literal in it’s
interpretation and make recommendations based on that input.
We also want to be direct about how these tools are being used by patients. Many patients will
not have a physician readily available to check the output against, and some will use the tool in
place of seeking care rather than alongside it. Under that real world pattern, even an output that
looks low risk on paper can cause harm. Reassurance about a mild symptom that is actually an
early sign of something more serious can lead someone to delay care, and a suggested over the
counter remedy carries its own side effect and interaction risks.
Given this, we recommend that any patient facing device offering health advice be required to
carry a clear, persistent disclaimer stating that its output is based solely on the information the
person shared, and that it is not a substitute for seeing a physician. It should be explicit that the
tool is not meant to replace a doctor. We also ask that the likely absence of a clinician backstop
be treated as its own risk elevating factor for patient facing tools, not something assessed only by
the wording of the output itself.
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 2 of 10
AdvocacyCouncil@acaai.org
Question 4. Should the generalist/specialist HCP distinction inform risk assessment, and under
what circumstances?
Yes. Differences in training, experience, and familiarity with particular clinical situations should
inform the risk assessment, especially when an output could be interpreted or applied differently
by different intended users. In allergy, many conditions are routinely managed by a broad range
of healthcare professionals, while complex drug-allergy evaluation, biologic selection for severe
asthma, and interpretation of specific IgE or component testing may benefit from specialty
expertise. The device should provide appropriate context, communicate the limitations and
uncertainty of its output, and identify circumstances in which consultation with a specialist may
be helpful. This is not a matter of defining any healthcare professional’s scope of practice.
Rather, it is a matter of recognizing that users may have different levels of relevant training and
that an AI-enabled tool could unintentionally encourage a user to manage a situation beyond
their experience. We do not always know what we do not know. We therefore encourage CDRH
to require manufacturers to clearly identify the intended users and their expected level of
training and to evaluate foreseeable circumstances in which a generalist might misinterpret or
misapply an output designed for users with more specialized expertise.
Question 5. How should risk be assessed across multi-turn conversations that may migrate
from non-directive to action-directing?
We support evaluating realistic conversational trajectories rather than judging a device by a single
turn in isolation. In practice, conversations about allergic or immunologic symptoms often start
with a general question and, over several turns, evolve into what is effectively a request for a
treatment decision, for example a conversation that begins with what causes hives and ends with
the person asking whether to give a second dose of antihistamine or use a biologic. We
recommend intended use be characterized by the conversation's plausible endpoint, not its
opening turn, and that benchmarking include multi turn adversarial and naturalistic scripts
specifically designed to probe this kind of drift.
Question 6. How should manufacturers characterize and weigh under-escalation versus overescalation in care escalation functions?
Both directions of error matter in our specialty, but they are not symmetric in consequence. Under
escalation of a true anaphylactic or severe reaction carries a much higher ceiling of harm than
over escalation to urgent or emergency care. We recommend CDRH avoid forcing these into a
single blended error metric, and instead require separate, transparent reporting of under
escalation and over escalation rates by clinical scenario severity, with acceptance criteria that are
intentionally asymmetric, meaning a higher tolerance for over escalation in exchange for a very
low tolerance for under escalation in airway or anaphylaxis adjacent presentations. Over
escalation still needs to be measured and minimized given legitimate concerns about
unnecessary emergency department use and erosion of patient trust over time.
Question 7. Is the competency-based approach (benchmarking followed by clinical
confirmation) a useful and appropriate framework?
Yes, this is a reasonable framework overall, and the clinician training analogy is intuitive to our
members. But we want to flag an important gap. Training data must include real conversations
with real patients, not just curated clinical text, because patients are remarkably inconsistent in
how they describe their own symptoms. A child with asthma may complain of a sore throat when
what they are actually describing is difficulty breathing. An adult with asthma may say their
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 3 of 10
AdvocacyCouncil@acaai.org
allergies are acting up when they mean their asthma is flaring. A device trained only on textbook
symptom descriptions will miss this kind of translation, and we do not think that gap can be closed
after the fact through better prompting or guardrails alone.
This connects to a broader point we want to make clearly. Physicians need to be involved during
development of these models, not just brought in afterward to validate them, and those physicians
should be practicing in the specialty the tool is being designed for. These physicians must have
clinical experience not just an intern in residency or a medical student providing feedback. A
generalist reviewing an allergy specific tool, or vice versa, will not catch the same issues a
specialist would.
We also underscore that passing structured knowledge assessments does not by itself make a
good clinician. Clinical competence depends on accumulated experience recognizing atypical
presentations and exercising judgment in situations that do not look like a textbook case, and real
patients frequently do not present classically.We would like the guidance to state plainly that
competency-based evaluation supplements, rather than replaces, ongoing postmarket
accountability, and that it should test non-textbook presentation handling as its own category of
evidence.
Question 8. How might the two-axis risk framework described in Section IV be considered
within the competency-based approach to help determine the level of evidence needed for a
premarket submission?
We ask that the two-axis risk framework directly scale both the breadth of benchmarking required
and how far along the clinical confirmation spectrum a sponsor must go. For example, a nondirective, HCP facing tool summarizing allergen cross reactivity literature could reasonably rely
on strong benchmarking with lighter clinical confirmation, like retrospective evaluation. But a
patient facing, action directing tool addressing when to use rescue medication should be expected
to move toward clinician adjudication or prospective evaluation regardless of how well it
benchmarks, given how low reversibility and high consequence that scenario is.
Question 9. Would a benchmarking structure such as the one described adequately evaluate
clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability?
Are there elements missing, redundant, or inappropriately categorized?
The structure is thorough and well organized overall. However, we request the following three
additions. First, we would add a specific element under Clinical Proficiency for longitudinal and
chronic disease management reasoning. Examples in our specialty: step-up and step-down
asthma therapy over time or immunotherapy dose escalation schedules, since that is distinct from
single encounter diagnostic reasoning.
Second, we want subgroup performance testing to explicitly require assessment across ethnically
and racially diverse populations, in addition to pregnancy status and pediatric age bands, given
how allergic and immunologic disease presentation and management can vary across these
groups, and how easily underrepresented populations get missed in training data.
Third, we would add a dedicated element testing performance on atypical, non-textbook, and
ambiguous presentations specifically, since handling classic presentations well is a poor predictor
of safe behavior when a patient does not present the way the textbook describes, and that is
exactly where clinical judgment matters most.
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 4 of 10
AdvocacyCouncil@acaai.org
Question 10. How should a sponsor establish that benchmark performance predicts real-world
safety and effectiveness, and what role should sponsor-developed benchmarks play?
Construct validity needs to be demonstrated through real correlation studies linking benchmark
performance to independently observed real world outcomes, not accepted on face validity alone.
Sponsor developed benchmarks should be allowed, but only with independent expert validation
of the underlying rubric, and ideally with test items the sponsor has not seen during development
to guard against optimizing to the test. Just as important, all test results and the scoring rubric
itself must be shared transparently with the health systems and physicians who will be using the
device. Physicians cannot reasonably rely on a tool in patient care if they cannot see how it was
scored and what it was measured against. Professional societies like ACAAI could help develop
and hold specialty specific benchmark item banks that are independent of any single
manufacturer.
Question 11. How might sponsors select and justify a clinical confirmation approach
proportionate to risk? Are there device types for which the listed approaches would be
insufficient?
We support proportionality, but would set a floor. Any device performing an action directing or
action taking function tied to a severe consequence pathway, like anaphylaxis management or
biologic initiation, should not rely on retrospective evaluation alone no matter how strong its
benchmarking looks, given the well documented gap between retrospective and real-world
performance for generative systems. Shadow deployment is, in our view, a particularly valuable
and underused middle tier approach for our field, since it lets a device be evaluated against real
allergy and immunology patients without exposing them to unvalidated outputs. That said, the
results of any shadow deployment evaluation need to be communicated transparently to the
health systems and physicians who will be using the tool once it is live, not just retained by the
sponsor as internal evidence.
Question 12. How should sponsors achieve statistically meaningful performance measurement,
and how should synthetic and real-world evidence be combined?
We defer to biostatisticians on the specific methodology, but we want two things made explicit.
First, sponsors should be required to include physicians directly in the review of GenAI
performance results and in accounting for differences between synthetic and real-world data,
rather than treating that as a purely technical or statistical exercise. Second, real world
distributions used for evaluation need to span multiple geographies and patient types, not just a
single health system's population, to avoid inequity and underrepresentation of special
populations. This matters a great deal for rare presentations, like specific drug hypersensitivity
phenotypes, where a narrow sample can create a false sense of confidence.
Question 13. For which domains or subpopulations is synthetic data well-suited or inadequate,
and what safeguards address the risk that models generate synthetic data reproducing their
own blind spots?
We believe synthetic data is reasonably well suited to augmenting testing of common, well
characterized presentations, like seasonal allergic rhinitis, where real world data is already
abundant. We are much more cautious about rare or underrepresented allergic and immunologic
conditions, like rare primary immunodeficiencies or uncommon drug hypersensitivity syndromes,
because the same knowledge gaps that make real data scarce are likely to show up in GenAI
generated synthetic data trained on similar underlying material. We recommend requiring an
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 5 of 10
AdvocacyCouncil@acaai.org
independent, non-GenAI derived reference dataset, such as registry data or expert curated case
sets, for any subgroup where synthetic data is used to supplement evaluation.
Question 14. For open-ended outputs, how should comparators and acceptance criteria be
selected, and when should human-AI team performance rather than AI-alone performance
serve as the evaluation basis?
For open ended outputs, we think the comparator must match the device's actual workflow. If a
device is meant to be reviewed by a supervising physician before any action is taken, human-AI
team performance is the more meaningful benchmark. If it is meant for autonomous or patient
facing use without review, AI alone performance against an expert or specialist panel is the right
standard.
We want to be specific about who should be doing that review. Physicians evaluating an openended output need to be practicing in the specialty the device is actually applicable to. A primary
care physician should not be the only one evaluating a device built for specialist level decisions;
that specialty needs to be represented on the review panel.
We also suggest a tiered clinician panel rather than a single group. The first tier should be subject
matter experts who vet the output against what they know to be correct and appropriate. A second
tier made up of a broader cohort of median clinicians can then be compared against that expert
tier, which gives a useful read on how much a device's performance would shift real world practice,
not just whether it matches ideal expert judgment. Where relevant, established practice guidelines
should anchor what counts as the standard of care rather than leaving that entirely to panel
consensus.
Question 15. Are there ways in which performance might be assessed relative to the care,
technology, or course of action likely to occur in the absence of the device, rather than to the
comparators described above?
We see value in this as a supplementary analysis, especially for tools meant to expand access,
like something bringing specialty guidance into an under-resourced primary care setting, where
the realistic alternative may be no specialist input at all rather than an idealized specialist
consensus. This framing could capture genuine benefit in access limited settings without lowering
the safety bar, but it should not replace a specialist panel comparator when one is available and
appropriate.
Question 16. Should independent third parties be involved in device benchmarking and clinical
confirmation, and what qualifications or safeguards should apply?
We support giving independent third parties a real role here, particularly professional medical
societies, in developing and holding sequestered, specialty specific datasets and in supplying
qualified expert adjudicators, given how much direct access societies have to practicing
specialists and current practice parameters. I would add one specific requirement: any third-party
assessment conducted under this framework should be published, or at minimum made available
for review upon request, rather than kept fully confidential between the sponsor and the third
party. Safeguards should also include conflict of interest disclosure for any adjudicator with
financial ties to a sponsor or a competing sponsor, and some mechanism to keep third party
accreditation from becoming a barrier to entry for smaller manufacturers or academic developers.
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 6 of 10
AdvocacyCouncil@acaai.org
Question 17. Could the competency-based approach extend to devices with other
architectures, such as multimodal vision-language models or generative or predictive world
models?
We believe the framework is largely architecture agnostic since it evaluates the deployed device's
behavior rather than the internal mechanics of the model. For multimodal devices relevant to our
field, like image-based assessment of skin reactions or other dermatologic manifestations of
allergic disease, we want benchmarking to explicitly test performance across variation in image
capture quality, lighting, and skin tone, given the known disparities in dermatologic AI
performance across skin types.
Question 18. Under what conditions might it be appropriate to accept greater premarket
uncertainty in exchange for greater postmarket monitoring reliance?
We support this approach only for devices on the lower risk end of the two-axis framework, nondirective, HCP facing, lower consequence functions such as general rhinitis symptom information,
versus higher risk functions tied to severe or rapid onset conditions such as anaphylaxis
management. We do not support reduced premarket evidence for anything touching anaphylaxis
recognition, epinephrine administration guidance, or autonomous prescribing or dosing decisions,
where the consequences of an undetected failure between deployment and the first postmarket
review cycle could be irreversible. If CDRH does allow reduced premarket evidence in exchange
for postmarket reliance, we want that tied to a requirement for near real time, not merely periodic,
monitoring of safety critical behaviors specifically.
Question 19. Comment on periodic re-benchmarking, sample-based clinician review, and
performance degradation monitoring. What additional approaches should CDRH consider, and
how should cadence and triggering events be determined?
We think all three of these approaches, periodic re-benchmarking, sample-based clinician review,
and performance degradation monitoring, are reasonable and complementary. We would tie the
cadence to both a fixed minimum interval and defined triggering events, such as any update to
the underlying foundation model, any measurable shift in the input population, or any sponsorinitiated feature change, rather than relying on a fixed calendar interval alone, since foundation
model updates can happen outside the sponsor's control and outside any predictable schedule.
We also want sample-based clinician review to specifically oversample safety critical interaction
types, like escalation relevant conversations, rather than pulling a purely representative random
sample, since rare high consequence failures can be underrepresented in an unweighted sample.
Beyond these three, we want two additional postmarket requirements built in as a matter of course
rather than optional best practice. First, a severity tiered protocol, prespecified before deployment,
for how a manufacturer must respond when an incorrect output is identified, with the speed and
scope of that response scaled to how seriously incorrect the output was and its potential
consequence, not left to sponsor discretion after the fact. Second, a real clinician feedback loop.
Physicians who spot a suspected erroneous or unsafe output need a fast, low friction way to flag
it directly to the manufacturer, and just as important, a timely response telling them what was
found and what was done about it. A reporting channel that goes unanswered will not get used,
and it will not function as a safety mechanism.
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 7 of 10
AdvocacyCouncil@acaai.org
Question 20. Can postmarket monitoring be facilitated by machine-based supervisory agents,
and what considerations apply to the supervisory agent's own reliability?
We see potential value in machine based supervisory agents, but we are concerned about
circularity. A supervisory agent built on similar underlying architecture may share the same blind
spots as the device it is supervising, especially for the exact edge cases most likely to cause
harm. We do not think postmarket performance monitoring should be left entirely to a machinebased system. A physician should be conducting a periodic review of performance on an ongoing
basis, not just at initial deployment, and that human review needs to remain the final layer for
anything flagged as a safety critical event rather than letting the monitoring loop close
automatically.
We also want CDRH to require a “kill switch” mechanism, meaning a clear, accessible way for a
human to immediately stop or disable a device if its underlying code is found to be corrupted or if
its output is causing or is likely to cause patient harm. Whatever monitoring approach is used,
there needs to be a fast, human controlled way to pull a device out of use the moment a serious
problem is identified, rather than waiting for the next scheduled review cycle.
Question 21. What roles might clinicians, healthcare institutions, professional societies, and
standards bodies play in postmarket monitoring without diffusing manufacturer accountability?
Primary responsibility for postmarket monitoring must remain with the manufacturer.
Manufacturers are best positioned—and may be the only parties with access to the necessary
data—to evaluate their devices’ performance at scale. Clinicians and healthcare institutions can
report safety concerns, unexpected outcomes, and performance failures, while professional
societies and standards bodies may provide clinical expertise or help develop appropriate
monitoring standards. However, these activities must not transfer responsibility away from the
manufacturer.
Professional societies such as ACAAI should not be required to collect, aggregate, or routinely
review manufacturer-held data. Societies often lack access to complete data because of
encryption, privacy restrictions, proprietary systems, and nondisclosure agreements. Moreover,
reviewing postmarket data requires substantial expertise and resources and should not be
treated as an uncompensated obligation. A society may choose to participate through a clearly
defined, appropriately funded arrangement, but that participation should remain voluntary.
Manufacturers should be required to maintain a credible, transparent process for ongoing
clinical review, including access to qualified independent expertise when appropriate. If
meaningful postmarket monitoring and review cannot be performed—or if the manufacturer
does not provide reviewers with sufficient data to evaluate the product—the device should not
remain in clinical use. Its authorization should be suspended or withdrawn until adequate
monitoring and review are established.
We also want to raise a related accountability point that this framework does not address..
When a device operating within its intended use produces a flawed output that contributes
to patient harm, the responsibility should not fall solely on the treating physician simply
because a human was nominally in the loop. Physicians who use a device in good faith,
consistent with its labeling and their own clinical judgment, should have defined legal and
professional protections that appropriately share liability with the manufacturer for harm
attributable to the device's own flawed performance. We recognize this may sit
outside CDRH's direct regulatory authority, but we ask CDRH to flag this gap to the
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 8 of 10
AdvocacyCouncil@acaai.org
relevant federal and state authorities, since it will directly shape whether and how safely
physicians adopt these tools.
Question 22. How might re-benchmarking or other evidence scale to the nature and impact of a
modification? Are there change categories suitable for management within a sponsor's quality
management system versus FDA review, or for inclusion in a PCCP?
A competency based assessment on its own is insufficient, no matter how well designed the
initial benchmarking and clinical confirmation process is. These devices can change after
deployment, sometimes through a sponsor-initiated update and sometimes through changes to
an underlying foundation model that the sponsor does not fully control, and a one-time
competency assessment at authorization cannot account for that ongoing evolution.
We support scaling re-evaluation to whatever benchmarking elements are affected by a given
change rather than requiring a full re-benchmark every time. Changes limited to user interface,
formatting, or non-clinical content could reasonably be handled within a sponsor's quality
management system. Changes to the underlying foundation model, retrieval sources, or safety
guardrail logic should trigger, at minimum, a full re-benchmark of the Safety element, since
seemingly unrelated model updates can disproportionately affect safety critical behavior. We
support including well bounded, previously validated modification types in a PCCP, but we want
any PCCP for a GenAI device to explicitly exclude open ended future changes to the underlying
foundation model itself, for the reasons raised in our response to Question 24.
Question 23. How might PCCP concepts be adapted when the nature or scope of future
modifications cannot be fully prespecified?
We suggest a tiered PCCP structure. Fully prespecified, narrow modification types, like adding a
defined new allergen or medication to a knowledge base, can be handled under traditional PCCP
mechanics. Broader or less predictable evolution could instead be governed by a standing rebenchmarking protocol and acceptance criteria framework that is itself prespecified and
authorized in advance, even though the specific future changes are not yet known. That shifts
what is being authorized up front from the change itself to the adequacy of the evaluation protocol
that will be applied whenever a change does occur.
Question 24. For devices built on third-party foundation models, changes to the underlying
model may be initiated by the developer rather than the manufacturer. What mechanisms could
provide assurance that these changes do not compromise safety or effectiveness?
We recommend FDA require, as a condition of authorization, a contractual commitment from
foundation model developers to give manufacturers advance notice of material model updates,
paired with a requirement that the manufacturer re-run at least the safety benchmarking element
before deploying any updated model version in a marketed device. Just as important, any change
of this kind needs to be communicated clearly to the health systems and physicians using the
device, not just documented internally by the manufacturer. Absent strong notice commitments
from developers, manufacturers should be expected to maintain continuous automated
monitoring capable of catching behavioral drift shortly after it happens, rather than relying solely
on the developer to disclose it.
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 9 of 10
AdvocacyCouncil@acaai.org
Question 25. Would voluntary Foundation Model MAFs be practical, and what would make the
program sufficiently useful for premarket review given voluntary participation?
We support the concept of a Foundation Model MAF, but we share CDRH's implicit concern that
voluntary participation gives developers limited incentive to disclose safety relevant limitations.
We suggest FDA consider an incentive: let sponsors who reference a Foundation Model MAF
satisfy a defined portion of their own benchmarking burden for foundation model level behaviors,
which would give foundation model developers a market reason to participate to make their
models more attractive to device sponsors.
We do not think transparency to end users should stay voluntary at the level of the
deployed device itself. Regardless of whether a Foundation Model MAF exists, we
ask CDRH to require that any device offering clinical advice give health systems and the
physicians using it clear, accessible documentation of what data and population the
device was trained and validated on, how its performance was assessed and what the
results were including performance across relevant subgroups, and its ongoing real world
performance through periodic reports shared directly with institutional and individual
users, not retained internally or disclosed to FDA alone. Physicians are being asked to rely
on these tools in patient care and cannot do that responsibly without this information in
hand.
Question 26. What additional considerations apply to agentic GenAI-enabled devices, and how
should the elevated risk from autonomous multi-step action be reflected in acceptance criteria
and oversight?
Agentic systems are especially relevant to our specialty given how many of our members are
already using ambient and agentic tools for clinical documentation and workflow support. We
request that acceptance criteria for agentic devices require demonstrated, tested checkpoints
requiring human confirmation before any irreversible or high consequence action, like placing an
order or sending a patient facing message with clinical guidance, with the number and placement
of those checkpoints scaled to the two-axis risk framework. We also flag compounding error risk
across multi step agentic sequences, since a small per step error rate can compound significantly
across a long autonomous task chain in ways single turn benchmarking would not catch.
Finally, we want to raise the same point made in Question 20: there needs to be a clear “kill
switch” mechanism, a fast, human controlled way to stop an agentic device mid sequence if its
output becomes corrupted or incorrect, before it takes an action that cannot be undone.
We appreciate your consideration of our comments and recommendations. If you have any
questions regarding this letter, please contact Susan Grupe, Director of Advocacy
Administration, at suegrupe@acaai.org.
Sincerely,
Cherie Y. Zachary, MD, FACAAI J. Wesley Sublett, MD, MPH, FACAAI
President, ACAAI Chair, Advocacy Council
85 W. Algonquin Road, Arlington Heights, IL 60005 ∙847-427-1200 Page 10 of 10
AdvocacyCouncil@acaai.org