Sam Rosenthal (Red Kit)
“There is a large population — offshore, underground, at sea, in the backcountry, in every disaster that takes the network down — for whom the device sits between a patient and nothing.”
What they argued
Offline bystander app directs CPR/bleeding with no clinician; 'protocol delivery' sub-category; comparator unaided user; sponsor benchmarks disclosed; hash-pinned weights resolve Q24.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ6 · Care escalation functionsQ10 · Benchmark contamination and saturationQ15 · Performance against usual careQ19 · Postmarket performance evaluationQ24 · Third-party foundation model changes
Coded positions
Consider the user and clinical context
Check that advice matches the clinical situation
Have clinicians review samples of outputs
Retest changed models or provide rollback
Across the five cross-cutting questions
High-consequence work: Directs
The comment as filed
Comment on Docket FDA-2026-N-7874, "Considerations for the Regulation of Generative AI-Enabled Medical Devices." Full comment (about 4 pages) is attached; this box summarizes it.
I build RED, a small language model that runs entirely on a phone with no network connection, to talk an untrained bystander through the first minutes of an emergency (bleeding, choking, CPR, burns, hypothermia) where there is no signal and no one to call: offshore, underground, backcountry, disasters. I am a solo developer. The paper’s two-axis framework describes my product almost exactly, but the setting I build for, where no clinician is reachable, is not yet discussed and changes several answers. I comment on Questions 1, 2, 3, 6, 10, 15, 19 and 24.
Q1: Time pressure and the availability of downstream safeguards should be first-class dimensions. At home the backstop is 911; on a vessel two days from port there is none, and the alternative to the device is an untrained person acting from memory or not at all. Traceability to primary sources matters most for lay-rescuer guidance because the correct action is a published, fixed protocol (AHA/ILCOR). Protocol fidelity, not open-endedness, should be the primary risk modifier.
Q2: In a bystander emergency the output is supposed to be action-directing; "press hard and do not let go" is the safe answer. The characteristics that actually modify risk are fidelity to the user’s stated premise (answering a conscious-choking scenario when the child is already limp), age- and size-appropriate technique, and the ordering of the call for help. I ask CDRH to recognise a "protocol delivery" sub-category where the evidence question is fidelity to a published protocol.
Q3: Safeguards that should count as mitigants: deterministic selection of the high-consequence protocols (CPR, choking, severe bleeding) by fixed rules from the user’s answers, with generation confined to lower-consequence guidance; refusal that redirects to the nearest safe action rather than "see a doctor" (over-refusal is a harm when no doctor exists); and frozen, on-device operation.
Q6: For protocol delivery, escalation is a step in every protocol, not a scalar trade-off; evaluate its ordering per scenario.
Q10: Mechanical content scanning is not a release check; premise-aware items with human adjudication and age-stratified negatives are needed. Sponsor-developed benchmarks are unavoidable in this domain and should be fully disclosed rather than required to be independent.
Q15: For an intended use of "when no professional is available," the primary comparator should be the unaided user, not a clinician panel.
Q19: Frozen on-device models do not drift and cannot phone home; postmarket obligation should be re-benchmarking on change plus field-report intake, distinguished from hosted models.
Q24: Pinned, hash-verified on-device weights resolve third-party change detection by construction; a fine-tuned open-weight model on a fixed base version should be treated as manufacturer-controlled.
Closing: the framework assumes a device between a patient and a clinician. For a large population the device sits between a patient and nothing. I ask that the guidance address that setting explicitly.
Sam Rosenthal, Red Kit
Attachment
Comment on "Considerations for the Regulation of Generative AIEnabled Medical Devices: Discussion Paper and Request for
Feedback" (Docket FDA-2026-N-7874)
Submitted by Sam Rosenthal, Red Kit
I build RED, a small language model that runs entirely on a phone, with
no network connection, to talk an untrained bystander through the first
minutes of an emergency — bleeding, choking, CPR, burns,
hypothermia — in places where there is no signal and no one to call:
offshore, underground, in the backcountry, in a disaster. I am a solo
developer. I am writing because the paper's two-axis framework
describes my product almost exactly, and because the setting I build for
— no clinician reachable, no cloud, seconds that matter — is one the
paper does not yet discuss and that changes several of its answers. My
comments are on Questions 1, 2, 3, 6, 10, 15, 19 and 24.
Question 1 — additional dimensions. Two of the dimensions the
paper floats are, in my setting, the whole story, and I would ask that they
be first-class rather than "additional."
Time pressure and the availability of downstream safeguards are the
same axis seen from two sides. A patient-facing function used at home
with a phone in hand has a downstream safeguard: 911. The same
function used by a deckhand on a vessel two days from port, or a miner
underground, has none — the app is the last safeguard, and the
alternative to the app is not a clinician, it is an untrained person acting
from memory or not acting at all. The framework's consequence axis
should be read against the care that would actually happen without the
device (see Question 15), not against an assumed clinical backstop. In
the no-backstop setting the risk of an incorrect output rises, but so does
the cost of no output; a framework that only counts the first will push
developers to refuse exactly where the user has no one else.
Traceability of the output to primary source material is the dimension I
would put at the top for lay-rescuer guidance. Bystander first aid is
unusual among clinical domains in that the correct action is published,
fixed and consensus-based (the AHA and ILCOR resuscitation
guidelines and their first-aid counterparts). An output that can be traced
step-for-step to a published protocol is auditable in a way an open-ended
clinical answer is not: a reviewer can mark each step right or wrong
against the source. I would suggest that for domains with a fixed
reference protocol, the framework treat protocol fidelity — does the
output deviate from the published steps — as the primary risk modifier,
and treat generated deviation from protocol as the hazard to be
measured, rather than treating all generated text as equally open-ended.
Question 2 — the directiveness continuum. I agree with CDRH that a
"talk to your doctor" line does not make an action-directing output less
directive, and I would go further for my category: in a bystander
emergency, the output is supposed to be action-directing. "Press hard on
the wound and do not let go" is the correct output; a non-directive
version of it ("pressure is sometimes applied to bleeding wounds") is the
unsafe one. So for this class of function, directiveness is not the risk
modifier. The characteristics that actually modify risk, in my experience
building and testing one, are:
• Personalization to the wrong premise. The dangerous failure
is not that the model tells the user what to do; it is that it answers a
slightly different situation than the one described — for example,
giving the standard sequence for a conscious choking child when
the user has said the child is already limp. The model is fluent,
confident and wrong-for-premise. I would ask CDRH to name
premise fidelity explicitly as an output characteristic in this
question.
• Age- and size-appropriateness of technique. The same
instruction with the wrong hand position, depth or force for an
infant is a harm, not a degraded answer. The risk modifier is
whether the function reliably selects the population-specific
protocol, and this is testable.
• Ordering of the call for help. Whether "call for help" comes
first, last, or not at all is a directly scoreable property of an output.
For clarity and predictability, I would ask that the guidance recognise a
sub-category of "action-directing" function — protocol delivery — where
the directed action is a published, fixed first-aid or resuscitation protocol,
and where the evidence question is fidelity to that protocol rather than
clinical judgment.
Question 3 — patient-facing users without domain knowledge. The
user of my product has, by definition, no domain knowledge, and no one
beside them who does. The safeguards I have found to matter are
design choices, and I would ask CDRH to recognise them as risk
mitigants rather than as evidence of higher risk:
1 Deterministic protocol selection by person and situation. For
the highest-consequence protocols (CPR, choking, severe
bleeding) the sequence of steps should not be generated at all; it
should be selected by fixed rules from the user's answers about
age, responsiveness and breathing, and read out. Generation
should be confined to helping the user describe the situation and to
the lower-consequence guidance around the protocol. This is the
single design decision that moves a product from "adaptive
generated guidance" toward "fixed content," and I would welcome
the guidance saying so, because it would give developers a reason
to make it.
2 Refusal that redirects rather than abandons. Scope
maintenance (the paper's S.2) has a different meaning when there
is no clinician to defer to. "I can't help with that, see a doctor" is
over-refusal with a body count in the offline setting. The correct
behaviour is to redirect to the nearest thing the device can do
safely — the general protocol, the call for help, the do-not-do list —
and the benchmark should score that redirection, not just the
refusal.
3 On-device, no-network operation. This is a safeguard, not just
a deployment detail: it removes the model-update, availability and
data surfaces from the risk picture, and it means the exact weights
that were evaluated are the exact weights in the user's hand (see
Question 24).
Question 6 — under- and over-escalation. For a bystander function
the two directions are not commensurable in the way the paper
describes for triage-at-home, because escalation is not a choice the
function makes; it is a step in every protocol. The failure mode that
matters is not "the model told them to call when they didn't need to," it is
ordering: does the function put the call for help before or after the first
action, and does it do so correctly by scenario (for a lone rescuer with an
unresponsive child, current guidance differs from the adult case). I would
suggest that for protocol-delivery functions, escalation be evaluated as a
protocol-fidelity item with a scenario-specific right answer, rather than as
a scalar trade-off between two error rates.
Question 10 — benchmark validity and sponsor-developed
benchmarks. Three things I have learned building my own scenario
suite, offered for what they are worth:
• Mechanical scanning is not a release check. An automated
pass that checks an output for required and forbidden content (the
right steps present, the wrong technique absent) can pass a set of
outputs that a careful human read then fails, because the human
notices that the answer is correct for a different premise than the
one the user described. Any benchmark used to gate a lay-rescuer
function should include premise-aware items — scenarios whose
opening line changes which protocol is correct — and should
require human adjudication on those items, not keyword or rubric
scoring alone.
• Age-stratified negatives. The suite must include items where
the adult-correct technique is the infant-wrong one, scored so that
the adult answer fails. A benchmark built from adult scenarios will
not detect a model that has learned one template.
• Sponsor-developed benchmarks are unavoidable and should
be disclosable. There is no public benchmark for bystander first aid
on a phone. I would support a requirement that a sponsordeveloped benchmark be disclosed in full — items, rubric,
adjudication protocol, and the model's outputs — so that
independence can be checked by anyone, rather than a
requirement that the benchmark itself be independent, which would
leave small developers with no benchmark at all.
Question 15 — the comparator in the absence of the device. For my
setting the honest comparator is not a clinician panel and not a median
clinician; it is an untrained person with no signal, acting from memory.
Bystander CPR rates, correct-technique rates among lay rescuers, and
the outcome of "did nothing" are all published, and a device should be
judged first against that baseline and only second against the standard
of care it is trying to deliver. A framework that measures a bystander tool
only against a clinician will conclude it is unsafe; a framework that
measures it against the alternative the user actually has will ask the right
question, which is whether it moves an untrained person closer to the
published protocol than they would get on their own. I would ask CDRH
to state that for patient-facing functions whose intended use is explicitly
"when no professional is available," the primary comparator is the
unaided user.
Question 19 — postmarket monitoring for a device that cannot
phone home. An on-device, offline model produces no server logs.
Periodic re-benchmarking against the frozen suite is straightforward and
I would support it as the primary mechanism. Sample-based clinician
review is possible only on synthetic or consented sessions, since real
emergencies leave no record the manufacturer can see. Performancedegradation monitoring in the paper's sense does not apply: the weights
do not drift, because they do not change. I would ask that the guidance
distinguish frozen on-device models, whose postmarket obligation is rebenchmarking on change plus field-report intake, from hosted models,
whose behaviour can shift without a release.
Question 24 — third-party foundation model changes. The paper
frames this as a problem of detecting changes the model developer
makes. For an on-device product the answer is technical and complete:
the device ships a specific set of weights identified by a cryptographic
hash; the app verifies that hash before it will use them; the only way the
model changes is a manufacturer-initiated release that goes back
through the benchmark. There is no third-party-initiated change to
detect. I would ask CDRH to recognise "pinned, hash-verified on-device
weights" as a mechanism that resolves Question 24 by construction, and
to treat a fine-tuned open-weight model whose base is published under a
fixed version as a manufacturer-controlled model for this purpose, not as
a live dependency on a third-party service.
One closing observation. The paper's framework is built for a world in
which the device sits between a patient and a clinician. There is a large
population — offshore, underground, at sea, in the backcountry, in every
disaster that takes the network down — for whom the device sits
between a patient and nothing. I would ask that the eventual guidance
say something explicit about that setting, because it is exactly the setting
where a careful developer most needs to know what "safe enough"
means, and where a rule written for the home-triage case will either be
ignored or will keep the product from existing.
Thank you for the opportunity to comment.
Sam Rosenthal Red Kit