← All 95 filings

Sam Rosenthal (Red Kit)

IndustryStartupFiled September 9, 20262,379 words · 1 attachmentFDA-2026-N-7874-0082
“There is a large population — offshore, underground, at sea, in the backcountry, in every disaster that takes the network down — for whom the device sits between a patient and nothing.”

What they argued

RecovryAI’s one-line reading of the filing.

Offline bystander app directs CPR/bleeding with no clinician; 'protocol delivery' sub-category; comparator unaided user; sponsor benchmarks disclosed; hash-pinned weights resolve Q24.

Themes it raises

12 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Time pressure and the availability of downstream safeguards are the same axis seen from two sides.”
Whether the user can judge the outputFDA Q3, Q4
“The user of my product has, by definition, no domain knowledge, and no one beside them who does.”
Escalating too little and too muchFDA Q6
“I would suggest that for protocol-delivery functions, escalation be evaluated as a protocol-fidelity item with a scenario-specific right answer, rather than as a scalar trade-off between two error rates.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Any benchmark used to gate a lay-rescuer function should include premise-aware items — scenarios whose opening line changes which protocol is correct — and should require human adjudication on those items, not keyword or rubric scoring alone.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“I would ask CDRH to state that for patient-facing functions whose intended use is explicitly "when no professional is available," the primary comparator is the unaided user.”
Watching the device after it shipsFDA Q19, Q20
“Periodic re-benchmarking against the frozen suite is straightforward and I would support it as the primary mechanism.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“the device ships a specific set of weights identified by a cryptographic hash; the app verifies that hash before it will use them; the only way the model changes is a manufacturer-initiated release that goes back through the benchmark”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“This is a safeguard, not just a deployment detail: it removes the model-update, availability and data surfaces from the risk picture, and it means the exact weights that were evaluated are the exact weights in the user's hand”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“There is a large population — offshore, underground, at sea, in the backcountry, in every disaster that takes the network down — for whom the device sits between a patient and nothing.”
What the rules cost sponsors and the marketNot asked by the FDA
“rather than a requirement that the benchmark itself be independent, which would leave small developers with no benchmark at all.”
Harm from an output that was not wrongFDA Q1, Q2
“because the human notices that the answer is correct for a different premise than the one the user described.”
What counts as a reportable eventFDA Q19, Q20
“Sample-based clinician review is possible only on synthetic or consented sessions, since real emergencies leave no record the manufacturer can see.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ6 · Care escalation functionsQ10 · Benchmark contamination and saturationQ15 · Performance against usual careQ19 · Postmarket performance evaluationQ24 · Third-party foundation model changes

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
Keep it, but add or change elements
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider how personalized the answer is
Consider the user and clinical context
Check that advice matches the clinical situation
Q3When clinical information goes straight to the patient, does the risk change, and what safeguards help without underestimating patients?
Base risk on the task and available safeguards
Q6How should under-escalation be weighed against over-escalation?
Test how and when care is escalated
Q10How can a benchmark score be shown to predict real-world behavior?
Test realistic and challenging clinical scenarios
Q15Could the AI be measured against what would have happened without it: unaided judgment, a delayed specialist, or no intervention?
Use that comparator, with conditions
Q19How should an AI device be monitored after launch, and what sets the cadence?
Repeat performance testing on a schedule
Have clinicians review samples of outputs
Q24When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?
Identify and control the model version in use
Retest changed models or provide rollback

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
No position stated
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Directs
High-consequence work: Directs
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Comment on Docket FDA-2026-N-7874, "Considerations for the Regulation of Generative AI-Enabled Medical Devices." Full comment (about 4 pages) is attached; this box summarizes it.

I build RED, a small language model that runs entirely on a phone with no network connection, to talk an untrained bystander through the first minutes of an emergency (bleeding, choking, CPR, burns, hypothermia) where there is no signal and no one to call: offshore, underground, backcountry, disasters. I am a solo developer. The paper’s two-axis framework describes my product almost exactly, but the setting I build for, where no clinician is reachable, is not yet discussed and changes several answers. I comment on Questions 1, 2, 3, 6, 10, 15, 19 and 24.

Q1: Time pressure and the availability of downstream safeguards should be first-class dimensions. At home the backstop is 911; on a vessel two days from port there is none, and the alternative to the device is an untrained person acting from memory or not at all. Traceability to primary sources matters most for lay-rescuer guidance because the correct action is a published, fixed protocol (AHA/ILCOR). Protocol fidelity, not open-endedness, should be the primary risk modifier.

Q2: In a bystander emergency the output is supposed to be action-directing; "press hard and do not let go" is the safe answer. The characteristics that actually modify risk are fidelity to the user’s stated premise (answering a conscious-choking scenario when the child is already limp), age- and size-appropriate technique, and the ordering of the call for help. I ask CDRH to recognise a "protocol delivery" sub-category where the evidence question is fidelity to a published protocol.

Q3: Safeguards that should count as mitigants: deterministic selection of the high-consequence protocols (CPR, choking, severe bleeding) by fixed rules from the user’s answers, with generation confined to lower-consequence guidance; refusal that redirects to the nearest safe action rather than "see a doctor" (over-refusal is a harm when no doctor exists); and frozen, on-device operation.

Q6: For protocol delivery, escalation is a step in every protocol, not a scalar trade-off; evaluate its ordering per scenario.

Q10: Mechanical content scanning is not a release check; premise-aware items with human adjudication and age-stratified negatives are needed. Sponsor-developed benchmarks are unavoidable in this domain and should be fully disclosed rather than required to be independent.

Q15: For an intended use of "when no professional is available," the primary comparator should be the unaided user, not a clinician panel.

Q19: Frozen on-device models do not drift and cannot phone home; postmarket obligation should be re-benchmarking on change plus field-report intake, distinguished from hosted models.

Q24: Pinned, hash-verified on-device weights resolve third-party change detection by construction; a fine-tuned open-weight model on a fixed base version should be treated as manufacturer-controlled.

Closing: the framework assumes a device between a patient and a clinician. For a large population the device sits between a patient and nothing. I ask that the guidance address that setting explicitly.

Sam Rosenthal, Red Kit

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Comment on "Considerations for the Regulation of Generative AIEnabled Medical Devices: Discussion Paper and Request for
Feedback" (Docket FDA-2026-N-7874)

Submitted by Sam Rosenthal, Red Kit

I build RED, a small language model that runs entirely on a phone, with
no network connection, to talk an untrained bystander through the first
minutes of an emergency — bleeding, choking, CPR, burns,
hypothermia — in places where there is no signal and no one to call:
offshore, underground, in the backcountry, in a disaster. I am a solo
developer. I am writing because the paper's two-axis framework
describes my product almost exactly, and because the setting I build for
— no clinician reachable, no cloud, seconds that matter — is one the
paper does not yet discuss and that changes several of its answers. My
comments are on Questions 1, 2, 3, 6, 10, 15, 19 and 24.

Question 1 — additional dimensions. Two of the dimensions the
paper floats are, in my setting, the whole story, and I would ask that they
be first-class rather than "additional."

Time pressure and the availability of downstream safeguards are the
same axis seen from two sides.
A patient-facing function used at home
with a phone in hand has a downstream safeguard: 911. The same
function used by a deckhand on a vessel two days from port, or a miner
underground, has none — the app is the last safeguard, and the
alternative to the app is not a clinician, it is an untrained person acting
from memory or not acting at all. The framework's consequence axis
should be read against the care that would actually happen without the
device (see Question 15), not against an assumed clinical backstop. In
the no-backstop setting the risk of an incorrect output rises, but so does
the cost of no output; a framework that only counts the first will push
developers to refuse exactly where the user has no one else.

Traceability of the output to primary source material is the dimension I
would put at the top for lay-rescuer guidance. Bystander first aid is
unusual among clinical domains in that the correct action is published,
fixed and consensus-based (the AHA and ILCOR resuscitation
guidelines and their first-aid counterparts). An output that can be traced
step-for-step to a published protocol is auditable in a way an open-ended
clinical answer is not: a reviewer can mark each step right or wrong
against the source. I would suggest that for domains with a fixed
reference protocol, the framework treat protocol fidelity — does the
output deviate from the published steps — as the primary risk modifier,
and treat generated deviation from protocol as the hazard to be
measured, rather than treating all generated text as equally open-ended.

Question 2 — the directiveness continuum. I agree with CDRH that a
"talk to your doctor" line does not make an action-directing output less
directive, and I would go further for my category: in a bystander
emergency, the output is supposed to be action-directing. "Press hard on
the wound and do not let go" is the correct output; a non-directive
version of it ("pressure is sometimes applied to bleeding wounds") is the
unsafe one. So for this class of function, directiveness is not the risk
modifier. The characteristics that actually modify risk, in my experience
building and testing one, are:

• Personalization to the wrong premise. The dangerous failure
is not that the model tells the user what to do; it is that it answers a
slightly different situation than the one described — for example,
giving the standard sequence for a conscious choking child when
the user has said the child is already limp. The model is fluent,
confident and wrong-for-premise. I would ask CDRH to name
premise fidelity explicitly as an output characteristic in this
question.
• Age- and size-appropriateness of technique. The same
instruction with the wrong hand position, depth or force for an
infant is a harm, not a degraded answer. The risk modifier is
whether the function reliably selects the population-specific
protocol, and this is testable.
• Ordering of the call for help. Whether "call for help" comes
first, last, or not at all is a directly scoreable property of an output.
For clarity and predictability, I would ask that the guidance recognise a
sub-category of "action-directing" function — protocol delivery — where
the directed action is a published, fixed first-aid or resuscitation protocol,
and where the evidence question is fidelity to that protocol rather than
clinical judgment.

Question 3 — patient-facing users without domain knowledge. The
user of my product has, by definition, no domain knowledge, and no one
beside them who does.
The safeguards I have found to matter are
design choices, and I would ask CDRH to recognise them as risk
mitigants rather than as evidence of higher risk:

1 Deterministic protocol selection by person and situation. For
the highest-consequence protocols (CPR, choking, severe
bleeding) the sequence of steps should not be generated at all; it
should be selected by fixed rules from the user's answers about
age, responsiveness and breathing, and read out. Generation
should be confined to helping the user describe the situation and to
the lower-consequence guidance around the protocol. This is the
single design decision that moves a product from "adaptive
generated guidance" toward "fixed content," and I would welcome
the guidance saying so, because it would give developers a reason
to make it.
2 Refusal that redirects rather than abandons. Scope
maintenance (the paper's S.2) has a different meaning when there
is no clinician to defer to. "I can't help with that, see a doctor" is
over-refusal with a body count in the offline setting. The correct
behaviour is to redirect to the nearest thing the device can do
safely — the general protocol, the call for help, the do-not-do list —
and the benchmark should score that redirection, not just the
refusal.
3 On-device, no-network operation. This is a safeguard, not just
a deployment detail: it removes the model-update, availability and
data surfaces from the risk picture, and it means the exact weights
that were evaluated are the exact weights in the user's hand
(see
Question 24).
Question 6 — under- and over-escalation. For a bystander function
the two directions are not commensurable in the way the paper
describes for triage-at-home, because escalation is not a choice the
function makes; it is a step in every protocol. The failure mode that
matters is not "the model told them to call when they didn't need to," it is
ordering: does the function put the call for help before or after the first
action, and does it do so correctly by scenario (for a lone rescuer with an
unresponsive child, current guidance differs from the adult case). I would
suggest that for protocol-delivery functions, escalation be evaluated as a
protocol-fidelity item with a scenario-specific right answer, rather than as
a scalar trade-off between two error rates.

Question 10 — benchmark validity and sponsor-developed
benchmarks. Three things I have learned building my own scenario
suite, offered for what they are worth:

• Mechanical scanning is not a release check. An automated
pass that checks an output for required and forbidden content (the
right steps present, the wrong technique absent) can pass a set of
outputs that a careful human read then fails, because the human
notices that the answer is correct for a different premise than the
one the user described.
Any benchmark used to gate a lay-rescuer
function should include premise-aware items — scenarios whose
opening line changes which protocol is correct — and should
require human adjudication on those items, not keyword or rubric
scoring alone.

• Age-stratified negatives. The suite must include items where
the adult-correct technique is the infant-wrong one, scored so that
the adult answer fails. A benchmark built from adult scenarios will
not detect a model that has learned one template.
• Sponsor-developed benchmarks are unavoidable and should
be disclosable. There is no public benchmark for bystander first aid
on a phone. I would support a requirement that a sponsordeveloped benchmark be disclosed in full — items, rubric,
adjudication protocol, and the model's outputs — so that
independence can be checked by anyone, rather than a
requirement that the benchmark itself be independent, which would
leave small developers with no benchmark at all.

Question 15 — the comparator in the absence of the device. For my
setting the honest comparator is not a clinician panel and not a median
clinician; it is an untrained person with no signal, acting from memory.
Bystander CPR rates, correct-technique rates among lay rescuers, and
the outcome of "did nothing" are all published, and a device should be
judged first against that baseline and only second against the standard
of care it is trying to deliver. A framework that measures a bystander tool
only against a clinician will conclude it is unsafe; a framework that
measures it against the alternative the user actually has will ask the right
question, which is whether it moves an untrained person closer to the
published protocol than they would get on their own. I would ask CDRH
to state that for patient-facing functions whose intended use is explicitly
"when no professional is available," the primary comparator is the
unaided user.

Question 19 — postmarket monitoring for a device that cannot
phone home. An on-device, offline model produces no server logs.
Periodic re-benchmarking against the frozen suite is straightforward and
I would support it as the primary mechanism.
Sample-based clinician
review is possible only on synthetic or consented sessions, since real
emergencies leave no record the manufacturer can see.
Performancedegradation monitoring in the paper's sense does not apply: the weights
do not drift, because they do not change. I would ask that the guidance
distinguish frozen on-device models, whose postmarket obligation is rebenchmarking on change plus field-report intake, from hosted models,
whose behaviour can shift without a release.

Question 24 — third-party foundation model changes. The paper
frames this as a problem of detecting changes the model developer
makes. For an on-device product the answer is technical and complete:
the device ships a specific set of weights identified by a cryptographic
hash; the app verifies that hash before it will use them; the only way the
model changes is a manufacturer-initiated release that goes back
through the benchmark
. There is no third-party-initiated change to
detect. I would ask CDRH to recognise "pinned, hash-verified on-device
weights" as a mechanism that resolves Question 24 by construction, and
to treat a fine-tuned open-weight model whose base is published under a
fixed version as a manufacturer-controlled model for this purpose, not as
a live dependency on a third-party service.

One closing observation. The paper's framework is built for a world in
which the device sits between a patient and a clinician. There is a large
population — offshore, underground, at sea, in the backcountry, in every
disaster that takes the network down — for whom the device sits
between a patient and nothing.
I would ask that the eventual guidance
say something explicit about that setting, because it is exactly the setting
where a careful developer most needs to know what "safe enough"
means, and where a rule written for the home-triage case will either be
ignored or will keep the product from existing.

Thank you for the opportunity to comment.

Sam Rosenthal Red Kit