Steven Zhao
Themes it raises
FDA questions it names
Q3 · When an output becomes directive
The comment as filed
See attached file(s)
Attachment
Supplemental Comment — Docket No. FDA-2026-N-7874 (Round 2: Agentic AI and
Foundation Models, **v1.0**)
Docket: FDA-2026-N-7874 — Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback (FDA / CDRH / DHCoE; released 2026-08-18) Comment period
closes: October 19, 2026 (verified 2026-09-26) Submitter type: Individual (no organizational affiliation) Submitter:
Steven Zhao — independent medical device regulatory practitioner Nature: Supplemental comment. Builds on my
earlier comment in this docket (posted 2026-09-15, docket comment id FDA-2026-N-7874-0093). This comment
does not repeat that submission; it addresses the paper's treatment of agentic AI systems and foundation models
with operational proposals.
Document note: this comment was drafted with AI tools and reviewed and approved by the named submitter, who
takes full responsibility for its content. The mechanisms described below were additionally prototyped in a small
working demonstration (a state-machine implementation, not a medical product) to verify they can be expressed in
runnable form; no claims of clinical validity are made or implied.
---
Executive summary
1. For agentic AI-enabled devices, I recommend FDA adopt an entity-level registration spine: the agent, the
model(s), the external tools/data sources it can invoke, and the datasets used to qualify it should each be registered,
versioned, and linked — so that "the device" is a reviewable configuration, not a moving target. 2. I recommend
framing agentic behavior change as evidence-gated state transitions: a change (model update, tool change, dataset
refresh, capability expansion) is a transition that is only valid when its pre-specified evidence requirements are met;
an unmet requirement makes the transition invalid rather than merely discouraged. 3. For postmarket monitoring, I
recommend FDA name silent failure patterns explicitly in its postmarket discussion: degraded input populations with
stable headline accuracy; tool or data sources that have aged out of validity without an error signal; and agent
behavior that deviates from declared workflows without raising exceptions. 4. I recommend the framework treat
human authorization as a required state transition for capability expansion: an agentic device should not be able to
move into a higher-impact capability tier through in-use learning or orchestration drift; that transition must be
human-authorized and evidence-carrying, with the authority to grant it kept distinct from the technical capability to
perform it. 5. These four recommendations are expressible with machinery manufacturers already run (design
controls, change management, audit trails) and are offered as operational additions to Sections VI–VII of the
discussion paper.
---
Thank you for the opportunity to provide a second round of input. I remain an independent practitioner submitting in
my personal capacity. My first comment responded to all 26 discussion questions with emphasis on benchmark
governance and change control. This supplemental comment responds specifically to the paper's discussions of
foundation models and agentic AI systems, because agentic deployments change the reviewable unit: what must be
evaluated and monitored is no longer only a model's outputs, but a configuration of entities whose composition can
change after clearance.
1. An entity-level registration spine for agentic devices
The discussion paper asks how risk assessment should account for agentic systems. I suggest the starting point is
definitional: an agentic device is not one artifact but a set of entities acting together — at minimum (i) the agent itself
(the orchestrating software entity that persists across sessions), (ii) the model or models it invokes, (iii) the external
tools and data sources it is authorized to call, and (iv) the datasets used to qualify the overall configuration.
For each entity class, registration would capture identity, version, and permissions (what that entity is allowed to do
within the device). Four practical consequences follow:
• Reviewable configuration. The premarket submission can pin a specific configuration graph ("agent v1.0 + model
v1.0 + imaging database v2026.09 + qualification dataset rev Q3"). Reviewers evaluate a frozen configuration,
exactly as they evaluate a frozen software version today.
• Traceable drift. Any postmarket deviation from the registered graph (a new tool invoked, a new model swapped in,
a dataset out of its declared validity window) becomes a detectable, reportable event rather than an invisible drift.
• Assignment of accountability. When a third-party foundation model sits behind the agent, the registration spine
makes explicit which entity the sponsor controls and which it only references — connecting to the paper's
Foundation Model Master Files discussion without conflating the sponsor's device with the foundation model.
• Audit artifacts. Each registration event and each configuration change leaves a hash-chained record; the audit trail
is the enforceable core, not an afterthought.
None of this requires new legal authority; it reuses device identification, software versioning, and supplier-control
concepts the industry already practices, extended to the agent/tool/dataset entities that the paper's current element
set touches only through the single agentic element (A.1).
2. Evidence-gated state transitions as the operational core of change control
My first comment proposed predetermined-gate change control as an evolution of PCCP. This comment proposes
the underlying formal shape, because agentic systems need it to be exact: every lifecycle-affecting change is a state
transition, and a transition is valid only when its declared evidence requirements are satisfied.
Concretely, a sponsor would declare, per transition type, the evidence that must exist before the transition may
occur. Illustrative minimum set:
Transition — Illustrative evidence requirement (pre-declared by sponsor) --- — --- Initial activation — Non-clinical
benchmark results; clinical confirmation per the paper's premarket approach Model update — Version manifest;
benchmark re-run; clinical review where risk class warrants Tool or data-source change — Provenance and validity
evidence for the new source; regression results Return to service after a monitoring pause — Re-verification
evidence; release note documenting the monitored condition and the recovery rationale
Two properties matter for regulation. First, invalidity, not discouragement: if the evidence requirement is unmet, the
transition does not occur — the system records a denial receipt. This gives FDA a binary, inspectable condition
instead of a judgment call made after deployment. Second, re-qualification is non-monotonic: a device that was
paused for a monitored condition returns to service only through a new evidence-carrying transition, not by the
passage of time or by a self-declared recovery. This "resume only through the gate" property is what keeps
postmarket pause authorities meaningful for self-adjusting systems.
The discussion paper's premarket competency frame evaluates whether the device performs as intended before
reaching patients. The state-transition frame extends the same evidentiary logic to every later change of the
deployed configuration, so that the premarket and postmarket halves of the framework share one spine rather than
two vocabularies.
3. Naming silent failures in postmarket monitoring
The paper invites comment on risk-proportionate postmarket monitoring. For agentic and continuously adjusting
devices, I recommend FDA explicitly name silent failure patterns — failure modes that produce no error signal and
can pass conventional performance dashboards:
• Population drift with stable headline accuracy. Aggregate accuracy can hold steady while performance on a shifted
subpopulation degrades. Monitoring should therefore track input-distribution drift as a first-class signal, not only
output quality.
• Expired inputs in active use. A tool, database, or dataset the agent still invokes may have moved outside its
declared validity window without generating an exception; staleness should be monitored as a lifecycle event, not
discovered at the next scheduled review.
• Path deviation without exceptions. An agent can complete tasks through workflows that diverge from its declared
operating envelope — no crash, no error log, but also no longer the behavior that was evaluated. Monitoring should
compare observed action trajectories against the declared envelope (echoing the bounded-conversation-envelope
construct in my first comment, generalized to non-conversational action).
For each pattern, the device's response should itself be a governed transition: pause, roll back to the last valid
configuration, re-verify, and resume only through the gate described in §2. Listing these patterns in guidance would
give sponsors a concrete monitoring checklist and give reviewers a common vocabulary for postmarket plans.
4. Human authorization as the gate for capability expansion
The paper asks how oversight should work for agentic systems. I recommend one bright-line rule: capability
expansion is a human-authorized transition. An agentic device should not acquire, through in-use learning,
orchestration drift, or tool composition, the ability to perform actions in a higher-impact tier (for example, from
retrieving information, to drafting records, to directly influencing clinical decisions) without an explicit,
evidence-carrying human authorization event. Two separations make this enforceable:
• Authority versus capability. The technical ability to perform an action must be kept distinct from the standing
authorization to perform it. A system may possess a capability that its current authorization tier does not permit;
invoking it should require the authorization transition, not merely the presence of the function.
• Pre-specification of tiers. The permission tiers (read-only retrieval; internal state changes; external record actions;
direct clinical influence) would be declared at design time, with the highest tiers default-denied and grantable only by
the human authority designated in the sponsor's quality system.
This keeps "machine requests, human decides" at exactly the points where error consequences step up in kind, not
merely in degree, and it is auditable: each tier crossing is a transition with a receipt.
5. Closing
The four recommendations above — an entity registration spine, evidence-gated state transitions, named
silent-failure patterns, and human-authorized capability expansion — are mutually reinforcing and deliberately
conservative: each reuses machinery (versioning, change management, audit trails, designated authorities) that
regulated industry already operates. My working demonstration of the state-transition and monitoring mechanics
suggests they are expressible in small, testable implementations; that prototyping experience, not any clinical claim,
is the basis for my confidence that these constructs are practical.
I appreciate the Agency's engagement on these questions and would be glad to see the agentic-AI portion of the
framework develop along the lifecycle-complete lines sketched above.
Submitted in my individual capacity, without organizational affiliation. Drafted with AI tools; reviewed and approved
by the named submitter, who takes full responsibility for the content.