Tianrui Huang
Themes it raises
FDA questions it names
Q2 · The spectrum of device activityQ4 · Generalist and specialist usersQ9 · The benchmarking structureQ12 · Statistically meaningful performanceQ17 · Devices with many functions
The comment as filed
I am a developer of AI-powered virtual reality (VR) software (software as an investigational device) for clinical exposure-based therapy delivery. The comments below are offered from my perspective and not as my company, which is a sponsor developing a generative AI-enabled VR platform used by clinicians to create exposure-based therapeutic content (e.g., objects, environments, AI-voiced avatars) for anxiety and OCD-spectrum disorders. I appreciate CDRH’s early engagement on this topic and offer the following responses to selected discussion questions where my experience raises considerations I believe are not yet fully captured in the paper.
Attachment
Comments on FDA-2026-N-7874: "Considerations for the Regulation of Generative
AI-Enabled Medical Devices: Discussion Paper and Request for Feedback"
Submitted by: Tianrui Huang, a developer of AI-powered virtual reality (VR) software (software
as an investigational device) for clinical exposure-based therapy delivery.
The comments below are offered from my perspective and not as my company, which is a
sponsor developing a generative AI-enabled VR platform used by clinicians to create
exposure-based therapeutic content (e.g., objects, environments, AI-voiced avatars) for anxiety
and OCD-spectrum disorders. I appreciate CDRH's early engagement on this topic and offer the
following responses to selected discussion questions where my experience raises
considerations I believe are not yet fully captured in the paper.
Response to Question 2 (directiveness of informational functions)
In my system, developed with clinicians, the clinician prompts the generative AI to create
specific exposure scenarios to trigger the patient before the patient is exposed to the generative
output of the AI. The information output is better characterized as clinician-directed content
generation and not a device-initiated recommendation to the patient. The clinician remains in full
control at all times as the point of clinical judgment, with the clinician deciding what the AI
generates based on the clinical prompt, and then the clinician reviews the generated content as
appropriate or not for the specific patient before it is shown to the patient.
I would suggest CDRH's framework to distinguish devices where the clinician is the recipient
and approver of a directive output from generative AI from devices where directive output
directly reaches the patient, without a clinician in between as the intervening clinical decision
point. The presence and timing of this clinician review step, where a clinician previews and
approves each generation before patient exposure, should sit lower on the activity axis than one
that generates and delivers generative output content autonomously.
Response to Question 4 (generalist vs. specialist HCP-facing)
Incorporating specialist-scope limitation into risk assessment for HCP-facing functions is
important, as devices intended for trained clinicians in a defined modality (such as exposure and
response prevention - ERP) rather than any licensed mental health clinician have a materially
different risk profile than a general-use clinical tool. The specialist’s training allows them to
supply contextual judgement (such as calibration of exposure intensity) that a device is not
independently responsible for. It would be important for CDRH to allow sponsors to define and
restrict user populations by modality-specific training or credentialing, with labeling and access
controls as a mitigation, as they can allow for the lowering of a device's risk tier, rather than
treating all HCP-facing functions as a single undifferentiated category.
Response to Question 9 (benchmarking structure adequacy)
The four proposed categories (Safety, Clinical Proficiency, Generalizability, Agentic) implicitly
assume that provoking distress, anxiety, or negative affect in the patient is itself a safety failure
to be minimized. But for device categories whose intended therapeutic mechanism is the
controlled generation of distress-inducing content, such as AI-generated exposure scenarios for
anxiety and OCD-spectrum disorders, this assumption doesn't hold.S.1 to S.3, as currently
framed, would misclassify correctly functioning devices as unsafe.
I would recommend either:
(A) A distinct benchmarking sub-element such as "therapeutic calibration," evaluating
whether generated content produces intensity appropriate to the prescribed exposure
hierarchy and patient tolerance, as opposed to either under-triggering (therapeutically
inert) or over-triggering (harmful, retraumatizing)
(B) Explicit guidance that S.1–S.3 be reinterpreted per-device based on the sponsor's
stated therapeutic mechanism, with acceptance criteria defined against the intended
intensity profile rather than a universal "avoid distress" baseline.
Response to Question 12 (combining benchmarking and clinical confirmation)
I would urge CDRH to add flexibility to this framework for small and early-stage sponsors, such
as academic-affiliated development teams, which may not be able to support separate
benchmarking and clinical confirmation evidence tracks.
A prospective feasibility-stage clinical study, using validated, standardized clinical outcome
measures appropriate to the therapeutic area, should, where the study design and endpoints
are prespecified with appropriate rigor, be permitted to serve double duty: generating both
clinical confirmation evidence and benchmarking-equivalent performance data. Requiring fully
independent evidence-generation efforts for both tracks risks concentrating generative AI
medical device development among only the largest, best-capitalized sponsors, and would
disadvantage precisely the academic and early-stage innovators who are often first to explore
novel clinical applications of this technology. A combined evidence pathway, scaled to the
device's risk tier, would better serve CDRH's stated least-burdensome principle without
compromising on evidentiary rigor.
Response to Question 17 (novel architectures, generalizability of competency-based
approach)
An important distinction for CDRH to note is generative AI systems that produce output (text,
image, audio, 3D/VR environments, avatar behavior) from predictive or generative world models
that simulate or forecast environmental state dynamics, commonly called World Models. These
AI models are architecturally and functionally distinct, as content-generation device central risks
are the appropriateness, calibration, and safety of what it produces in response to a clinical
prompt, evaluated primarily through the Safety and Clinical Proficiency elements described in
Section V.B. For a world-model device, the central risk surface is the fidelity and reliability of its
predictive/simulated dynamics over time, which is evaluated more likely through the
Generalizability and Robustness elements (R.1–R.2) and may warrant benchmarking
approaches closer to those used for predictive AI-enabled devices generally.
Treating both under a single "generative AI" umbrella risks either over-specifying evaluation
requirements for content-generation devices (importing world-model-style dynamic validity
testing where it doesn't apply) or under-specifying them for world models (importing
content-appropriateness benchmarks that don't capture predictive fidelity risk). The
competency-based framework would benefit from including an explicit branching point
distinguishing these two categories before benchmarking elements are selected.