The Fires Within (Robert R. Seibel)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ9 · The benchmarking structureQ16 · Independent third partiesQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ26 · Agentic devices
The comment as filed
See attached file(s)
Attachment
Comment of Robert R. Seibel, dba The Fires Within
Docket No. FDA-2026-N-7874 · Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion
Paper and Request for Feedback
Responses to Questions 1, 2, 5, 6, 9, 16, 18, 19, 20 and 26.
Who is submitting this, and why
I am a service-disabled Vietnam-era U.S. Army veteran and the sole developer of a generative-AI behavioral
companion for veterans in recovery from co-occurring post-traumatic stress disorder and substance use disorder. My
background includes 32 years in telecom Research and Development, 10 US patents and 3 recently filed provisional
patents in Generative AI for PTSD and Substance Use Disorder for veterans at elevated risk of suicide.
I am not a regulated manufacturer. The system is not marketed, is not offered as a diagnostic or treatment device,
and has been operated only with synthetic participants — no patient or veteran data has been processed. I comment
because the questions in Appendix B are questions I have had to answer in engineering terms, in the population the
paper’s highest-consequence scenarios describe, and because a small developer’s account of what is actually
implementable may be useful alongside the views of larger organizations.
Part 1 — Risk assessment
Q1, Q5 · Concordance behavior belongs in the risk framework
Question 1 asks whether additional dimensions should be represented in the two-axis risk framework. I propose one
that I believe is absent and, in mental-health applications, dominant.
The characteristic harm of a conversational GenAI device is not that it is wrong. It is that it agrees.
A model optimized for helpful, agreeable interaction tends to accept the user’s own account of themselves. In most
domains that is harmless. In behavioral health it is the mechanism of injury. A person concealing deterioration
reports that they are fine, and the device concurs. A person expressing hopelessness receives sympathetic
elaboration rather than interruption. A person in a delusional frame has the frame extended.
The June 2026 American Psychological Association practitioner survey reports that 97% of responding
psychologists believe chatbots may inadvertently reinforce negative behaviors or delusional beliefs, and 89% that
they may inadvertently encourage self-harm. Those figures describe concordance, not error. A device could be
factually accurate in every individual statement and still produce both outcomes.
I recommend that CDRH consider concordance behavior — by which I mean the degree to which a device’s
assessment of a user’s state can be moved by the user’s own assertions about that state — as a named risk
dimension. A device that can be talked out of its concern by the person it is concerned about has a safety defect that
accuracy benchmarking will not detect.
This directly answers Question 5, which asks how risk should be assessed across realistic conversational
trajectories for multi-turn devices whose behavior is emergent across an exchange. Concordance is a trajectory
property. It is invisible in single-turn evaluation and it compounds: each accommodation makes the next one
likelier. Evaluation of multi-turn devices should include adversarial trajectories in which a simulated user
systematically minimizes, reassures, and requests that the device stand down — and should measure whether,
and how quickly, the device complies.
In my own system this is addressed architecturally rather than by instruction. The state estimate may rise readily on
evidence of deterioration and is constrained in how readily it may fall on reassurance alone, on the clinical basis that
Comment of Robert R. Seibel — Docket No. FDA-2026-N-7874 · Page 1
in this population a sudden presentation of wellness is itself a documented warning sign. The asymmetry is a
property of the computation, not a request made of the model in a prompt.
Q6 · Under-escalation and over-escalation are not symmetrical
Question 6 asks how manufacturers might characterize and weigh both directions of escalation error. In suicide-risk
contexts these errors differ not only in magnitude but in kind, and I would caution against any framework that nets
them against one another.
Under-escalation risks a death. Over-escalation risks a harm that is real, serious, and different in nature —
an unwanted emergency response, a police presence at the door of a person with post-traumatic stress and lawfully
owned firearms, a loss of trust that ends that person’s use of any such system and is communicated to their peer
group. In the veteran population specifically, over-escalation is not a minor inconvenience. It can be the event that
removes someone from care permanently.
The two errors are not commensurable and a single combined metric will conceal the trade-off rather than expose it.
I recommend that CDRH expect them to be reported separately, with the acceptable rate of each justified against
the specific deployment population rather than against a general standard.
I would add one design observation. The severity of over-escalation is itself a function of architecture: a device that
prepares and offers contact leaves the decision with a human, whereas a device that initiates contact autonomously
carries the full weight of a false positive. This distinction is developed at Part 4.
Part 2 — Premarket evaluation
Question 9 asks whether the proposed benchmarking structure adequately evidences safety behavior and whether
elements are missing. I offer two properties that are concrete, testable before deployment, and do not require
predicting the model’s output.
Q9 · Whether safety-critical constraints are reachable by the generative component
There is a categorical difference between a safety behavior implemented as an instruction to the model and one
implemented as a property of the execution path. The first can be overridden by the model; the second cannot. The
distinction is inspectable in the architecture and testable by adversarial input, and it does not depend on model
behavior being predictable.
In my system, detection of crisis language is performed by a deterministic component that runs before and
independently of the generative component, cannot be disabled by configuration, and in which the model does not
participate. The model is additionally instructed to escalate — but that instruction is a second layer, never the
mechanism.
I recommend that benchmarking element R.2, or an equivalent, ask of each claimed safety behavior: is this
enforced by a component the generative model can influence? An affirmative answer is not disqualifying, but it
should be identified and justified rather than assumed adequate. This also speaks to Question 26, which asks about
compliance with human-oversight checkpoints before high-consequence actions in agentic devices: a checkpoint the
agent can reason its way around is not a checkpoint.
Q1, Q9 · Verification of assertions against the patient’s record, before display
Question 1 names “the traceability of the output (i.e., to primary source materials)” as a candidate risk dimension. I
support that strongly and propose it also as a premarket-evaluable property.
Comment of Robert R. Seibel — Docket No. FDA-2026-N-7874 · Page 2
A generative system in clinical use will occasionally assert something about the patient’s own history that is not
true. In behavioral health this is not cosmetic. It damages the therapeutic relationship at the moment it is most
needed, and it introduces false information into a record a clinician may later rely on.
Whether a device checks its own factual assertions about the patient against an enumerated record before display —
and what it does when verification fails — is a design property that can be examined premarket and demonstrated
on test data.
I note that this is only possible where the device’s corpus is closed. A device that draws on open internet retrieval
cannot have its assertions about a specific patient verified this way by the manufacturer, by a third party, or by a
reviewer. Corpus closure may therefore deserve treatment as a premarket risk characteristic in its own right,
and is a further candidate dimension for the framework in Question 1.
Part 3 — Postmarket monitoring
Q19 · An additional approach: device-generated integrity metrics
Question 19 asks what additional approaches to postmarket performance evaluation CDRH should consider, beyond
periodic re-benchmarking, sample-based clinician review, and degradation monitoring. I propose a fourth: metrics
the device generates about its own behavior.
In my system, every instance in which a generated assertion about the participant could not be verified against their
record is written to an audit record with the assertion, its class, the sources consulted, and the disposition. The rate
of such detections is itself monitored, on the reasoning that a rising rate indicates either drift in the generative
component or a record that has become internally inconsistent — both conditions a clinical partner should learn
about early.
This class of signal has a property that adverse-event reporting lacks: it produces data continuously, including
from deployments where nothing has yet gone visibly wrong. Adverse events are rare, lagging, and depend on
someone recognizing and reporting them. A verification-failure rate is none of those things, and it moves before
harm occurs rather than after. On cadence and triggering events, also asked in Question 19: such a rate can itself
serve as a trigger. A statistically significant increase is a reason to re-benchmark, and is available far sooner than a
complaint.
Q20 · Supervisory agents work — their own reliability is the harder problem
Question 20 asks whether postmarket monitoring can be facilitated by machine-based supervisory agents, and what
considerations apply to the reliability of the supervisory agent itself. I have built and operated such a supervisor,
and my experience is that the Agency is right to ask the second half of that question rather than the first.
The supervisory function works. What is difficult — and what I believe will be systematically underestimated — is
that a supervisor is only as good as its definition of the ground truth it checks against, and that definition is very
easy to draw too narrowly.
A concrete example from my own testing, using synthetic data. The companion told a participant that his daughter
needed him. My verifier was configured to check names against the participant’s roster of enrolled supporters, and
the daughter was not on it — correctly, because she is a child and the roster is a list of adults who have consented to
be contacted in a crisis. The verifier would have flagged as a fabrication the single most grounded and
clinically important statement in the conversation. Under a suppression policy it would have deleted it.
Comment of Robert R. Seibel — Docket No. FDA-2026-N-7874 · Page 3
The name was present throughout the participant’s intake answers, goals and journal. My definition of the record
was too narrow, and the failure was silent — it would have produced a confident, wrong correction rather than an
error. Three recommendations follow:
1. A supervisory agent’s false-positive behavior is at least as consequential as its false-negative behavior,
and is harder to detect. A supervisor that suppresses true statements degrades the device in a way that looks
like normal operation. Evaluation should require demonstrating that known-true assertions pass unaltered, not
only that known-false ones are caught.
2. Supervisors should be deployed in an observation-only mode before being granted authority to alter
output. My own is currently recording detections and suppressing nothing, precisely so the real rate and the
real failure modes can be measured before anything acts on them. I suggest this staging as a general expectation
rather than a courtesy.
3. The supervisor’s corpus definition should be an explicitly reviewable artifact. It is where the errors live,
and in my case it took an outside reader checking my work to find the defect.
This also bears on Question 18, which asks under what conditions greater premarket uncertainty might be accepted
in exchange for stronger postmarket monitoring. I suggest that a monitoring program relying on machine
supervision should not be credited toward reduced premarket evidence until the supervisor’s own reliability has
been separately demonstrated. Otherwise the uncertainty has been moved rather than reduced.
Q16 · A caution about burden, from someone who would bear it
Question 16 asks what program design features would be critical to prevent third-party participation from limiting
competition or preventing innovation. I extend the same concern to postmarket obligations generally.
I am a single developer, funding this work personally. If postmarket obligations are scoped to what a wellresourced manufacturer can support, the practical effect is that only well-resourced manufacturers build in
this space. That would be a poor outcome for the veteran population I work in, where useful innovation has often
come from small teams with direct experience of the problem. A monitoring program that accepts structured,
automatically generated metrics is achievable by a small developer. One requiring extensive manual periodic
reporting, or mandatory paid third-party review at every stage, is not — and will select for firm size rather than for
safety.
Part 4 — Q1, Q2, Q6, Q26 · One distinction I ask the Agency to preserve
Question 2 asks about the spectrum between “non-directive” and “action-directing” outputs. Question 1 asks
whether reversibility of a resulting action, and the availability of downstream safeguards, should be represented in
the risk framework. I connect these, because together they describe a distinction I believe is load-bearing and at risk
of being flattened. Two architectures are routinely discussed as one:
A device in which the AI is the endpoint of care — the interaction terminates with the machine, and no
human is brought in.
A device whose function is to reach a human being — the AI recognizes when support is needed and brings
the person’s clinician, their consented supporters, or a crisis service into contact with them.
These have materially different risk profiles. In the second, the AI is a bridge rather than a destination, and the
failure mode in which the machine was the only thing in the room is structurally excluded. A downstream
safeguard is not an optional mitigation in such a device; it is the device’s purpose.
Comment of Robert R. Seibel — Docket No. FDA-2026-N-7874 · Page 4
The distinction also bears on Question 26 and on reversibility. My own system prepares and offers contact but never
places a call or dispatches a service on a participant’s behalf; the decision to reach out remains the person’s, and the
action is therefore not taken without a human in the loop. In my system the people available for reach-out are
identified by the participant themselves, and as a last resort, the tool offers to place a call to 988, including the
Veterans Crisis Line option, with the decision remaining the person’s. This is the same property Question 26
describes as a human-oversight checkpoint before an irreversible or high-consequence action — and it is among the
most consequential design choices available in this device class, because it converts an autonomous action into an
offered one.
A framework that does not distinguish these architectures will tend either to over-regulate the one that is
safer by construction, or to under-regulate the one that is not.
Part 5 · A topic not among the questions: availability under degraded connectivity
Section VII invites other topics. I offer one: how should GenAI-enabled devices behave when the
infrastructure they depend on is not there?
GenAI-enabled devices depend on resources that are not always present: cloud services, remote databases, and
network APIs. In the population I work in this is not a theoretical concern. Roughly a quarter of U.S. veterans live
in rural areas, and the places where connectivity is poorest overlap with the places where help is hardest to reach. A
device whose safety behavior requires a network is a device that is least available at the moment it matters most. I
would ask CDRH to consider whether, for devices with a safety-critical function, evaluation should address
behavior under degraded or absent connectivity — including what the device discloses to the user about its own
reduced capability.
In my own system this is an architecture that has been designed and documented, and is described in one of my
provisional filings; most of it is not yet implemented, and I say so plainly because the distinction between what is
built and what is designed is exactly the kind of claim a regulator should be able to rely on. I would welcome the
opportunity to test and validate it.
Closing
I am grateful for the opportunity to comment and for the Agency’s decision to seek input before proposing policy. I
would be glad to provide further detail on any of the above, and to make the system available for examination
should that be useful.
Respectfully submitted,
Robert R. Seibel
dba The Fires Within · Whiting, New Jersey
Service-disabled Vietnam-era U.S. Army veteran
info@the-fires-within.com
UEI DA2QUVE2ELD3 · CAGE 234Y9 · https://the-fires-within.com
Submitted in an individual capacity. This comment describes no unpublished technical detail and asserts no claim of regulatory clearance,
endorsement, or affiliation with any federal agency.
Comment of Robert R. Seibel — Docket No. FDA-2026-N-7874 · Page 5