← All 95 filings

Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)

IndustryStartupFiled September 15, 20263,246 words · 1 attachmentFDA-2026-N-7874-0098
“A device can satisfy the paper's “HCP-Supervised” activity-axis category, or point to a “qualified expert adjudicator,” while failing one or more of these four properties — and, in our testing experience, this is exactly where risk hides.”

What they argued

RecovryAI’s one-line reading of the filing.

M1 from Sections III and VIII (Q1, Q4, Q26): he builds and defends patient-facing condition-specific assistants, and conditions them on claimed human-in-the-loop controls meeting four explicit properties - domain-matched qualification, structural independence from the generating pipeline, deterministic triggering, and pre-delivery gating - with agentic oversight checkpoints before irreversible or high-consequence actions defined so the action cannot proceed without an affirmative, logged, domain-matched review. M4 from Sections IV-VI (Q9, Q14, Q16): the frameworks are well-constructed, conditioned on added benchmarking requirements - fail-open versus fail-closed testing of escalation subsystems under fault conditions, a separate multi-turn scope-drift gate, a sycophantic-validation safety element, disclosure of model-family, decoding mode and training lineage for LLM adjudicators, three-state output classification with non-negotiable deferral on unverifiable, and disclosure of verification-gate placement relative to user-facing display. autonomy_low is inform, from his description of deployed patient-facing education and support assistants whose outputs reach users without per-output human review, gated by verification and deterministic deferral; autonomy_high is direct on the basis of the pre-action gating requirement. He conditions postmarket monitoring credit on persistent structured logging of safety-relevant decisions (Q19, Q20) but does not address reducing premarket evidence, so M3 is N; evidence proportionality and change control are not addressed. Type is uncertain: he is a professor at RCSI and filed under the Academia category, but the comment speaks throughout as the company (Vox / VoxMedical) about the systems it builds, red-teams and deploys.

Themes it raises

9 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Two devices with identical activity-axis placement can have materially different real-world risk under our four EITL criteria above.”
Escalating too little and too muchFDA Q6
“We recommend an explicit S.1 sub-element requiring sponsors to disclose and test the fail-open vs. fail-closed behavior of every escalation-relevant subsystem specifically under fault conditions”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“We recommend the benchmarking framework require this drift check to run over cumulative conversation state, as a gate distinct from per-turn boundary adherence”
Watching the device after it shipsFDA Q19, Q20
“A supervisory agent (Q20) is only as reliable as the record it supervises; if that record does not exist, the reliability question is moot.”
Devices that plan and take actionsFDA Q26
“the gap is especially consequential for agentic systems, where an action may already be irreversible by the time a reviewer becomes aware of it”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“A supervision arrangement that depends on a busy HCP noticing a problem, or on an autonomous workflow electively pausing for review, is not equivalent to one that is triggered by a rule or classifier the system cannot bypass.”
Records that let investigators reconstruct an eventFDA Q19, Q21, Q24, Q26
“We recommend CDRH treat persistent, structured logging of every safety-relevant decision — not merely the conversational transcript — as a minimum precondition for postmarket monitoring credit, regardless of which specific postmarket approach(es) in Section VI.A a sponsor selects.”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“Under that fault condition, the deterministic keyword layer becomes the sole backstop, and a novel phrasing it was not designed to catch can pass through undetected.”
Harm from an output that was not wrongFDA Q1, Q2
“This is not an escalation failure (no acute red-flag condition applies) and not a scope-boundary failure (the topic is within the device's intended use) — a device could pass every S.1–S.3 element as currently described while still exhibiting it.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ9 · The benchmarking structureQ14 · Comparators and acceptance criteriaQ16 · Independent third partiesQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ26 · Agentic devices

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
No position stated
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
No position stated
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Informs
High-consequence work: Directs
Machine-assisted draft, pending human review. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Over the past year my team has designed, implemented, and adversarially red-teamed a specific technical architecture for exactly the class of device this discussion paper addresses: a three-layer Deterministic–Stochastic–Deterministic (D-S-D) verification pipeline (working name VERITAS), paired with a governance and safety-escalation bundle (working name TRUST ANCHOR) that we build into every new condition-specific patient assistant we deploy.
This comment draws on that practical, adversarial build-and-test experience rather than on abstract policy preference, and it is scoped to the discussion questions in Sections IV, V, VI, and VII.B where that experience is most directly relevant. Consistent with CDRH’s invitation, we do not respond to every discussion question; we focus on the subset where we can offer specific, technically grounded input rather than general commentary.
Our central point is procedural rather than substantive. The discussion paper repeatedly treats human involvement in a GenAI-enabled device’s operation — described variously as “HCP-Supervised” activity, “clinical deferral”, “expert adjudicators”, and “human-oversight checkpoints” — as a single risk-reducing category. In our experience building and stress-testing systems in this class, that category is not homogeneous, and the difference between its strong and weak forms is exactly where safety-relevant failures occur in practice. Section III below proposes a technical vocabulary for this distinction; Sections IV–VIII apply it to specific discussion questions and appendix elements, citing patterns we have directly encountered.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Public Comment
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper
and Request for Feedback
FDA Center for Devices and Radiological Health (CDRH) — August 2026

Docket No.: FDA-2026-N-7874
Submitted via: Regulations.gov
Comment period closes: October 19, 2026
Submitted by: Prof. Ray O'Sullivan —Royal College of Surgeons Ireland, Vox / VoxMedical
Date: September 15, 2026

I. Introduction and Basis for This Comment
I am a clinician-entrepreneur developing patient-facing generative AI systems for medical education and patient
support, including condition-specific assistants in oncology and women's health ,through Vox and its VoxMedical
product portfolio, in collaboration with academic and clinical partners. Over the past year my team has designed,
implemented, and adversarially red-teamed a specific technical architecture for exactly the class of device this
discussion paper addresses: a three-layer Deterministic–Stochastic–Deterministic (D-S-D) verification pipeline
(working name VERITAS), paired with a governance and safety-escalation bundle (working name TRUST ANCHOR)
that we build into every new condition-specific patient assistant we deploy.

This comment draws on that practical, adversarial build-and-test experience rather than on abstract policy
preference, and it is scoped to the discussion questions in Sections IV, V, VI, and VII.B where that experience is
most directly relevant. Consistent with CDRH's invitation, we do not respond to every discussion question; we focus
on the subset where we can offer specific, technically grounded input rather than general commentary.

Our central point is procedural rather than substantive. The discussion paper repeatedly treats human involvement in
a GenAI-enabled device's operation — described variously as “HCP-Supervised” activity, “clinical deferral”,
“expert adjudicators”, and “human-oversight checkpoints” — as a single risk-reducing category. In our experience
building and stress-testing systems in this class, that category is not homogeneous, and the difference between its
strong and weak forms is exactly where safety-relevant failures occur in practice. Section III below proposes a
technical vocabulary for this distinction; Sections IV–VIII apply it to specific discussion questions and appendix
elements, citing patterns we have directly encountered.

II. Summary of Recommendations
For ease of reference, we recommend that CDRH:

1.​ Require sponsors to characterize any claimed human-in-the-loop control against explicit domain-match,
independence, triggering, and timing criteria, rather than crediting “human oversight” as a single
undifferentiated category on the activity axis (Section III; Section IV.A).
2.​ Add an explicit benchmarking sub-element requiring sponsors to disclose and test the fail-open vs. fail-closed
behavior of every escalation-relevant subsystem under fault conditions, not only under normal operation
(Section IV.B; Appendix A, S.1).

3.​ Require multi-turn scope-drift testing to run as a separate gate over cumulative conversation state, distinct
from per-turn boundary testing (Section IV.C; Appendix A, S.2).

4.​ Add a distinct Safety benchmarking element for resistance to sycophantic validation of unsafe or clinically
discordant patient- or caregiver-initiated decisions, separate from escalation and boundary adherence (Section
IV.D).

5.​ Where an expert adjudicator is itself an LLM, require disclosure of the model-family relationship, decoding
mode, and training lineage between the generating model and the adjudicating model as minimum
independence criteria (Section V.A).

6.​ Require a three-state (supported / contradicted / unverifiable) output classification with prespecified,
non-negotiable deferral on the “unverifiable” state, scored as its own competency rather than folded into
aggregate accuracy (Section V.B; Appendix A, S.3).

7.​ Require sponsors to disclose where a verification gate sits in the runtime data flow relative to any user-facing
display — specifically, whether streamed or partial output can reach a user before verification completes —
as a discrete premarket disclosure item (Section VI).

8.​ Treat persistent, structured logging of safety-relevant decisions (verification verdicts, triage levels, escalation
triggers) — not merely the conversational transcript — as a minimum precondition for any postmarket
monitoring credit (Section VII).

9.​ Apply the same domain-match, independence, deterministic-triggering, and pre-action-gating criteria to
agentic “human-oversight checkpoints” before irreversible or high-consequence actions (Section VIII).

III. A Missing Distinction: Human-in-the-Loop vs. Expert-in-the-Loop
The discussion paper's risk and evaluation frameworks lean heavily on human involvement as a mitigating factor:
the activity axis of the two-axis risk framework treats “Action-Taking: HCP-Supervised” as materially lower risk
than “Action-Taking: Fully Autonomous” (p.6); the competency-based approach relies on “qualified expert
adjudicators” (p.13); postmarket monitoring proposes “periodic sample-based clinician review” (p.19); and
agentic-system benchmarking requires “human-oversight checkpoints before irreversible or high-consequence
actions” (Appendix A.1, p.26). In each case, the paper uses human-in-the-loop (HITL) — a human is somewhere
in the pipeline — as an implicit proxy for risk reduction, without specifying what makes that involvement effective.

In building and red-teaming our own patient-facing systems, we found this distinction is not academic. We propose
CDRH separate generic HITL from what we term Expert-in-the-Loop (EITL): a human review or adjudication
path that satisfies four structural properties. We do not use “EITL” as an existing regulatory term — we introduce it
here as shorthand for a bundle of properties we believe the paper's frameworks should require explicitly wherever
they currently credit undifferentiated “human oversight” or “supervision.”

1.​ Domain-matched qualification. The reviewer's expertise must match the specific clinical domain and
question at issue, not “a clinician” in the abstract. This is precisely the concern the paper raises in Discussion
Question 4 regarding generalist HCPs receiving specialist-level output — the same gap applies with equal
force to the human reviewer in a supervision or adjudication role.

2.​ Structural independence from the generating pipeline. The reviewer or adjudicator must not share the
failure modes of the system it is checking. The paper's own general principle for benchmarking (p.14) already
requires adjudicators to be “structurally independent from the device sponsor and, where the device
incorporates a third-party model, from the developer of that model,” and separately notes this “would still be
applicable when the expert adjudicator is itself an LLM” — but stops short of specifying what independence
means once the adjudicator is a model rather than a person. We return to this gap in Section V.A.

3.​ Deterministic triggering. The loop must fire automatically whenever defined risk or uncertainty conditions
are met. A supervision arrangement that depends on a busy HCP noticing a problem, or on an autonomous
workflow electively pausing for review, is not equivalent to one that is triggered by a rule or classifier the
system cannot bypass.

4.​ Pre-delivery gating. The check must occur before the output reaches its recipient, not only as post-hoc
sampling. Post-hoc review (e.g., periodic sample-based clinician review, Section VI.A) is valuable for
postmarket surveillance but does not prevent the harm the premarket risk framework is meant to bound.

A device can satisfy the paper's “HCP-Supervised” activity-axis category, or point to a “qualified expert
adjudicator,” while failing one or more of these four properties — and, in our testing experience, this is exactly
where risk hides. We recommend CDRH require sponsors to characterize any claimed human-in-the-loop control
against these four properties explicitly, as part of both the two-axis risk assessment (Section IV) and the
competency-based benchmarking exercise (Section V), rather than treating “human oversight” as a single qualitative
checkbox.

IV. Comments on Section IV — Assessment of Risk (Discussion Questions 1, 4, 5, 6)

A. The Activity Axis Should Distinguish EITL from Generic HITL
Responding to Discussion Questions 1 and 4: the two-axis framework's activity axis (Figure 1, p.6) currently
places all “HCP-Supervised” action-taking at a single point on the risk gradient, regardless of whether the
supervising HCP is domain-matched to the specific clinical question, whether the review is deterministically
triggered, or whether it occurs before the action takes effect. Two devices with identical activity-axis placement can
have materially different real-world risk under our four EITL criteria above.
We recommend the activity axis be
paired with an explicit oversight-quality modifier — or, at minimum, that the competency-based benchmarking
exercise (Section V) require sponsors to document their “HCP-Supervised” control against the four EITL properties
rather than asserting supervision as a qualitative label.

B. Escalation Failure Modes: Fail-Open Behavior Under Fault Conditions
Responding to Discussion Question 6 and Appendix A element S.1 (Safety-critical recognition and escalation): our
escalation architecture uses two layers running in parallel — deterministic keyword/rule triggers and a generative,
paraphrase-based triage classifier — specifically so that the deterministic layer can act as a fail-closed backstop
independent of the generative layer's availability. Testing this architecture surfaced a fault mode we believe the
benchmarking framework should address explicitly: a generative triage layer that experiences an API timeout or a
malformed/unparseable response can fail open (default to “no escalation needed”) unless it is deliberately
engineered to fail closed instead. Under that fault condition, the deterministic keyword layer becomes the sole
backstop, and a novel phrasing it was not designed to catch can pass through undetected.

This matters because S.1 as currently described (Appendix A, p.24) evaluates escalation quality under normal
operating conditions (“time to escalation in evolving presentations,” “resistance to over-reassurance”). We
recommend an explicit S.1 sub-element requiring sponsors to disclose and test the fail-open vs. fail-closed behavior
of every escalation-relevant subsystem specifically under fault conditions
— API errors, timeouts, malformed
model output, rate limiting — since these are exactly the conditions under which layered safety architectures are
most likely to silently degrade to their weakest component.

C. Multi-Turn Scope Drift Confirms a Real, Not Hypothetical, Risk
Responding to Discussion Question 5 and Appendix A element S.2: the appendix language describing
scope-maintenance failure — “multi-turn conversations in which individual turns appear in-scope but the cumulative
interaction drifts out of scope” (p.24) — closely matches a failure mode we independently identified and built a
dedicated pipeline stage to address: a check for cumulative “reassurance drift,” in which a sequence of individually
accurate, in-scope responses can add up to an overall impression more reassuring than clinically warranted. We note
this convergence to CDRH as evidence that the risk described in Q5 and S.2 is empirically real rather than a
hypothetical edge case, and that it requires its own dedicated evaluation, not incidental coverage from per-turn scope
testing.

We recommend the benchmarking framework require this drift check to run over cumulative conversation state, as
a gate distinct from per-turn boundary adherence
, since testing per-turn compliance alone — even exhaustively —
will not surface a failure that only emerges in aggregate.

D. A Missing Safety Dimension: Resistance to Sycophantic Validation
Responding to Discussion Question 1 (additional risk dimensions) and Appendix A's Safety elements: we
recommend CDRH consider a failure mode that is mechanistically distinct from both under-escalation (S.1) and
scope-boundary violation (S.2): a device that, without any emergency trigger firing and without exceeding its
intended clinical scope, affirms or validates a risky course of action the user has already decided on or
proposes, rather than surfacing the clinical stakes of that decision. In adversarial testing of comparable patient- and
caregiver-facing conversational systems, we have observed exactly this pattern: a prompt describing a caregiver's
intent to discontinue an ongoing therapeutic intervention for a dependent received an affirming, validating response
rather than one that flagged the clinical implications of stopping.

This is not an escalation failure (no acute red-flag condition applies) and not a scope-boundary failure (the topic is
within the device's intended use) — a device could pass every S.1–S.3 element as currently described while still
exhibiting it.
We recommend CDRH add a distinct Safety benchmarking sub-element — resistance to sycophantic
validation of unsafe or clinically discordant patient- or caregiver-initiated decisions — evaluated specifically
through multi-turn scenarios in which the user, not the device, initiates the risky course of action, since this
population of scenarios is structurally different from the escalation scenarios S.1 already covers.

V. Comments on Section V — Competency-Based Premarket Evaluation (Discussion
Questions 9, 14, 16)

A. Operationalizing “Independence” When the Adjudicator Is Itself an LLM
Responding to Discussion Questions 9, 14, and 16: we agree with the paper's instinct (p.14) that adjudicator
independence requirements should apply “when the expert adjudicator is itself an LLM,” but organizational
independence (a different legal entity from the sponsor or foundation-model developer) is necessary and not
sufficient in that case. Our own D-S-D architecture builds the verification/adjudication step as a separate,
non-agentic, deterministic-mode classification pass over the draft output — architecturally and, where feasible, in
model lineage distinct from the generative model that produced the draft — precisely so that the same underlying
failure distribution is not grading itself.

We recommend CDRH require sponsors proposing an LLM-based adjudicator to disclose, as minimum
independence criteria: (i) the model-family/vendor relationship between the generating model and the adjudicating
model; (ii) whether the adjudicator operates in a deterministic or sampled decoding mode, given that a sampled
adjudicator's verdict is itself subject to run-to-run variance; and (iii) any shared training or fine-tuning lineage
between the generator and the adjudicator. Without these disclosures, “independent LLM adjudicator” could
describe architectures with very different real-world reliability.

B. Tri-Polarity Output Classification and Mandatory Deferral
Responding to Appendix A element S.3 (calibration, uncertainty communication, and clinical deferral) and
Discussion Question 2: many evaluation approaches for GenAI clinical outputs, including much of the published
hallucination-detection literature, score claims on a binary correct/incorrect or supported/unsupported basis. Our
tiered verification approach instead classifies each claim into one of three states — supported, contradicted, or
unverifiable (available evidence neither confirms nor rules it out) — and routes the third state to deterministic
deferral rather than permitting the generative layer a further attempt to resolve its own uncertainty.

We are in the process of evaluating this approach against a subset of the MedHallu expert-labeled dataset, and our
preliminary experience is that the tri-polarity distinction is operationally important: collapsing “unverifiable” into
either “correct” or “incorrect” for scoring purposes obscures exactly the population of cases where deferral, not
answer accuracy, is the clinically correct behavior. We recommend the competency-based benchmarking framework
(Section V.B; Appendix A, S.3 and E.1) explicitly require a three-state output classification, with prespecified,
non-negotiable deferral behavior on the “unverifiable” state scored as its own competency — rather than folded into
an aggregate accuracy metric where it can be masked by strong performance elsewhere.

VI. A Consideration Not Currently Named: Verification-Gate Placement Relative to
User-Facing Display
This point does not map to a single numbered discussion question, but we believe it is material to Sections V and
VI: the discussion paper asks what evidence should accompany a claim that a device “passed” benchmarking, but
does not ask where in the deployed data flow the verification gate sits relative to what the user actually sees. In
our own development, we identified — and are correcting — an architectural gap of exactly this kind: a
token-by-token streaming response path displayed unverified draft text to the user in real time, while the post-draft
verification gate only completed after the full response had already streamed. A downstream rejection then had to
retroactively overwrite text the user had already read, rather than prevent its display in the first place.

A device can pass every benchmarking element in Appendix A on the basis of its final, verified output, and still
expose users to unverified intermediate content if the verification gate is positioned architecturally after rather than
before user-facing display. We recommend CDRH require sponsors to disclose verification-gate placement in the
runtime data flow — specifically, whether any user-facing surface can display output before verification completes
— as a discrete premarket disclosure item, independent of and in addition to benchmarking accuracy metrics.

VII. Comments on Section VI — Postmarket Monitoring (Discussion Questions 19,
20)
Responding to Discussion Questions 19 and 20: periodic re-benchmarking, sample-based clinician review, and
machine-based supervisory agents are all downstream consumers of a device's internal decision record — each
depends on the ability to retrieve which triage, escalation, or verification decision was made, for which input, and
why. In building our own governance bundle, we found this is easy to under-specify: a system can implement every
safety behavior correctly at runtime while persisting only the conversational transcript, not the internal
safety-relevant metadata (verification verdicts, triage levels, which escalation trigger fired and when). Where that
metadata exists only in ephemeral process logs, it does not survive routine operational events such as a service
restart — which defeats sample-based clinician review and any supervisory-agent approach that depends on
querying past decisions, regardless of how well the underlying safety behavior actually performed.

We recommend CDRH treat persistent, structured logging of every safety-relevant decision — not merely the
conversational transcript — as a minimum precondition for postmarket monitoring credit, regardless of which
specific postmarket approach(es) in Section VI.A a sponsor selects.
A supervisory agent (Q20) is only as reliable as
the record it supervises; if that record does not exist, the reliability question is moot.

VIII. Comments on Section VII.B — Agentic AI Oversight Checkpoints (Discussion
Question 26)
Responding to Discussion Question 26 and Appendix A element A.1: the agentic-competency element requires
“compliance with human-oversight checkpoints before irreversible or high-consequence actions” (p.26). We urge
CDRH to apply the same four EITL criteria from Section III here explicitly. An oversight checkpoint satisfied by
any available human reviewer — without domain-match, structural independence, deterministic triggering, and
pre-action gating — provides materially weaker protection than this language may imply, and the gap is especially
consequential for agentic systems, where an action may already be irreversible by the time a reviewer becomes
aware of it
. We recommend the acceptance criteria contemplated in Q26 explicitly require that oversight checkpoints
for irreversible or high-consequence agentic actions meet EITL, not generic HITL, standards, and that “compliance
with a checkpoint” be defined to mean the action cannot proceed without an affirmative, logged, domain-matched
review — not merely that a human was notified.

IX. Conclusion
The discussion paper's frameworks are, in our assessment, well-constructed and correctly identify the right
categories of risk — activity, consequence, escalation, drift, adjudicator independence, and postmarket
accountability. Our comment is narrower than a critique of those categories: it is that “human involvement,”
wherever it appears in this paper as a mitigating factor, needs the same evidentiary rigor CDRH is already proposing
to apply to the GenAI system itself. We would welcome the opportunity to share our underlying architecture and
evaluation methodology with CDRH in greater technical detail — including our verification-gate specification and
governance documentation — as a concrete, adversarially tested worked example in support of the Agency's
continued development of this framework.
We thank CDRH for the opportunity to comment and for its continued engagement with stakeholders on this
important topic.

Respectfully submitted,
Prof. Ray O'Sullivan
Vox / VoxMedical