← All 95 filings

Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)

IndustryStartupFiled September 15, 20263,682 words · 1 attachmentFDA-2026-N-7874-0100
“Where these safeguards are present and demonstrated, patient-facing delivery should not, by itself, elevate a function's risk classification.”

What they argued

RecovryAI’s one-line reading of the filing.

M1 from Q3: opposition to a categorical presumption that patient-facing functions are higher risk, conditioned on demonstrated safeguards - outputs anchored to validated protocols with traceability, deterministic escalation safety nets that cannot be suppressed by conversational drift, structured deferral pathways to human care, and comprehension validation across health-literacy levels; Q2 adds that for the highest-consequence presentations greater directiveness is often the safer design. M2 from Q1 and Q7/Q8: evidence should scale with grid position rather than technology label. M4 from Q7/Q8 and Q14: support for benchmarking followed by clinical confirmation evaluated on the deployed device, with the median qualified clinician as default comparator, human-AI team performance where the workflow is human-in-the-loop, and counterfactual comparators permitted with justification. M3 from Q18: strong support for the premarket-postmarket trade, conditioned on a prespecified monitoring plan with sampling frames, cadences, analyses and thresholds, defined triggering events, independent clinician adjudication and defined corrective action, and stated to be less appropriate for fully autonomous high-consequence action-taking. M5 from Q24 and Q25: architectural containment of safety-critical logic outside the third-party model, version pinning with contractual notice, shadow-mode gating with a regression battery as promotion gate, and foundation model version changes that pass that battery within a defined envelope included in a PCCP. autonomy_low is act from their Q26 position that agentic care coordination, documentation, outreach and workflow support deliver autonomy gains at low clinical risk and should stay outside device oversight, plus their deployed scheduling automation; autonomy_high is direct because they require non-bypassable human-oversight checkpoints before irreversible or high-consequence actions. Type is arguable: a founder-led venture-stage company that nonetheless describes a production platform deployed across millions of patient interactions and a Defense Health Agency contract, so large is defensible.

Themes it raises

16 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“We recommend traceability be treated as a formal, verifiable mitigating factor on the consequences axis.”
Whether the user can judge the outputFDA Q3, Q4
“A categorical presumption that patient-facing functions are higher risk would penalize exactly the tools most capable of closing access gaps, and would embed the medical parentalism the paper cautions against.”
Escalating too little and too muchFDA Q6
“The two error directions are not commensurable: the marginal harm of a missed emergent presentation is categorically different from the marginal harm of an unnecessary urgent-care visit, and the appropriate trade-off varies by chief complaint, population, and care context.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“Clearstep supports the competency-based structure of non-clinical device benchmarking followed by clinical confirmation, evaluated on the final user-facing device as configured for deployment.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“In our experience, the strongest evidence that a benchmark predicts real-world behavior is a demonstrated linkage between benchmark performance and independently adjudicated real-world performance in the same deployed configuration.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“Default to the median qualified clinician in practice, not an idealized consensus panel.”
Trading premarket certainty for postmarket monitoringFDA Q18
“Clearstep strongly supports accepting greater premarket uncertainty in exchange for robust postmarket monitoring, for a reason the paper identifies: for open-ended conversational devices, no feasible premarket test can enumerate the input space, while real-world deployment generates exactly the evidence that premarket testing cannot.”
Watching the device after it shipsFDA Q19, Q20
“Sampling should be stratified by acuity and by interaction pattern, not uniform, because safety signal concentrates in high-acuity and atypical-trajectory interactions”
Who is accountable when something goes wrongFDA Q21
“a deploying healthcare institution’s participation in postmarket monitoring, local validation, or configuration of a device within the manufacturer’s prespecified, validated envelope should not, by itself, expose that institution to classification as a manufacturer or specification developer”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“Sponsors should pin model versions and secure contractual minimum-notice windows for deprecations, with continued access to prior versions during revalidation.”
Devices that plan and take actionsFDA Q26
“Characterized failure envelopes for tool errors and unavailability, demonstrating that the system degrades to a safe state (deferral or human handoff) rather than improvising around a failed dependency.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“Mandatory human-oversight checkpoints before irreversible or high-consequence actions, verified as non-bypassable in the deployed configuration rather than merely documented”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“Prompt-injection resistance evaluated across all input surfaces (user input, retrieved content including EHR-derived content, and tool outputs), since agentic devices embedded in clinical systems ingest substantial third-party content”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“A categorical presumption that patient-facing functions are higher risk would penalize exactly the tools most capable of closing access gaps, and would embed the medical parentalism the paper cautions against.”
How this fits rules that already existFDA Q8, Q9, Q16, Q25
“We support the paper’s recognition that agentic systems performing care coordination, clinical documentation, patient outreach, and workflow support may not be the focus of device oversight, and we urge CDRH to preserve that boundary clearly”
What the rules cost sponsors and the marketNot asked by the FDA
“Multiple accredited bodies per clinical domain, with published accreditation criteria and transparent, capped fee schedules, so certification does not become a fixed-cost moat favoring incumbents”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
Supports with conditions
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Acts
High-consequence work: Directs
Machine-assisted draft, pending human review. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Comment from Clearstep

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Clearstep Inc.
Docket No. FDA-2026-N-7874
[DATE], 2026

Dockets Management Staff (HFA-305)
Food and Drug Administration
5630 Fishers Lane, Rm. 1061
Rockville, MD 20852

RE: Docket No. FDA-2026-N-7874, Considerations for the Regulation of Generative AI-Enabled Medical
Devices: Discussion Paper and Request for Feedback

Dear Dockets Management Staff:

On behalf of Clearstep Inc. (Clearstep), thank you for the opportunity to provide comments in response
to the discussion paper issued by the Center for Devices and Radiological Health (CDRH) and its Digital
Health Center of Excellence, “Considerations for the Regulation of Generative AI-Enabled Medical
Devices.” Clearstep commends CDRH for seeking early stakeholder input before proposing policy, and
we support the paper’s central orientation: a risk-proportionate, least-burdensome, total product life
cycle approach in which evidentiary expectations scale with what a function does and the consequences
of relying on an incorrect output, not with the mere presence of generative AI.

Clearstep operates one of the most widely deployed AI-driven care navigation and triage platforms in
U.S. health systems, spanning patient-facing digital self-triage, voice automation for nurse advice lines
and contact centers, and clinician-facing intake and decision support. We therefore have direct
operating experience with nearly the full span of the two-axis risk framework the paper proposes, and
with the premarket and postmarket evidence questions it raises. In summary, our comments emphasize:

• Risk framework: We support the two-axis framework and recommend two additional riskmodifying dimensions: traceability of outputs to primary clinical source material, and reversibility
of the resulting action. Both are properties manufacturers can engineer for and FDA can verify.
• Care escalation: Under- and over-escalation should both be assessed, but they are not
commensurable; manufacturers should prespecify and justify a clinically grounded escalation
operating point, with performance reported in both directions and stratified by acuity.
• Architecture-aware evaluation: The competency-based approach should recognize architectural
risk controls. Hybrid systems that anchor safety-critical behaviors (escalation, refusal, disposition)
in deterministic, expert-encoded clinical logic can guarantee reproducibility properties that purely
generative systems can only estimate statistically, and benchmarking burden should reflect that
difference.
• Comparators: For open-ended outputs, the default comparator should reflect the care actually
available to the intended user: the median clinician in practice and, where justified, the
counterfactual of no device. Human-AI team performance should serve as the evaluation basis
wherever the deployed workflow is human-in-the-loop.
• Postmarket rebalancing: We strongly support accepting greater premarket uncertainty in
exchange for robust, prespecified postmarket monitoring, and we describe a sample-based
independent clinician review program of the kind CDRH proposes that Clearstep already operates
in production at scale.
• Foundation model change control: Manufacturers can and should be accountable for third-party
model changes through architectural containment, version pinning with contractual notice,
shadow-mode evaluation, and regression benchmarking as a promotion gate, mechanisms well
suited to Predetermined Change Control Plans (PCCPs).

About Clearstep Inc.
Clearstep is a leader in AI-driven healthcare access and navigation solutions, offering one of the most
prevalently implemented AI self-triage systems on U.S. health system websites, as documented in a
recent study in Nature’s npj Digital Medicine. Clearstep also serves as the routing engine for the Defense
Health Agency’s Digital Front Door+, supporting a diverse beneficiary population in a large, integrated
delivery environment.

Our platform is built on gold-standard clinical guidelines co-authored by Dr. Barton Schmitt, a co-creator
of the Schmitt-Thompson telephone triage protocols used in the majority of U.S. nurse call centers. We
integrate expert-encoded systems with statistical methods, including machine learning and large
language models (LLMs), in a hybrid architecture: deterministic, protocol-anchored clinical logic governs
safety-critical behaviors, while generative components support communication, information elicitation,
and documentation. Our predictions are designed to be explainable, interpretable, and tunable, with
every triage or navigation recommendation traceable back to validated clinical best practices and
subject to independent clinician review. Clearstep’s deployments include:

• AI Triage & Navigation tools that guide individuals to appropriate levels and modes of care.
• AI Voice Automation that augments nurse advice lines and call centers by handling common call
reasons and improving first-contact resolution.
• AI Optimized Scheduling that matches patient needs to available resources, integrated with major
EHRs via standards-based APIs where feasible.
Because Clearstep’s products operate across the activity spectrum, from non-directive information
through action-directing patient-facing triage to supervised voice automation embedded in clinical
workflows, and are evaluated continuously in real-world deployment, our comments below draw on
operating evidence rather than hypotheticals. Comments are organized by the sections and numbered
discussion questions of the paper.

I. Considerations for the Assessment of Risk (Section IV; Questions 1, 2, 3, 6)
A. The two-axis framework, and two additional dimensions (Question 1)
Clearstep supports the two-axis framework. Device activity and consequence-of-reliance are the correct
primary dimensions, and the framework’s implication, that evidence should scale with grid position
rather than with technology label, is essential to least-burdensome regulation. We recommend that
CDRH formally incorporate two of the candidate dimensions named in Question 1 as risk modifiers:

• Traceability to primary source material. An output whose basis can be independently inspected
(for example, a triage disposition traceable to a specific, validated clinical protocol and decision
pathway, or a summary linked to the underlying chart excerpts) permits meaningful independent
review by clinicians and local governance bodies, and materially reduces the consequence of an
incorrect output because errors are detectable before reliance. Traceability is an engineerable
property: manufacturers can design for it, sponsors can document it, and FDA can verify it. This
consideration also aligns with the independent-review criterion referenced in the paper’s
discussion of clinical decision support. We recommend traceability be treated as a formal,
verifiable mitigating factor on the consequences axis.

• Reversibility of the resulting action. A function whose downstream action is easily reversed or
intercepted (e.g., a routing suggestion reviewed by scheduling staff) presents a different risk than
one whose consequences are immediate and difficult to unwind. Reversibility interacts naturally
with the activity axis and would sharpen the distinction between supervised and autonomous
action-taking.

B. The directiveness continuum (Question 2)
We agree that directiveness is a continuum determined by the substance and context of an output, not
solely by trigger words, and we agree that boilerplate statements such as “talk to your doctor” should
not, by themselves, render an action-directing output non-directive. To give manufacturers clarity and
predictability, we recommend that CDRH:

• Publish worked examples along the continuum for common function types (the paper’s four-step
lisinopril illustration is a useful template), so manufacturers can deliberately design and document
where each output class sits;
• Assess directiveness at the level of the function’s characterized output classes (e.g., a defined
disposition set) rather than sentence-by-sentence, so that a bounded, prespecified set of
dispositions, each mapped to validated clinical guidance, can be evaluated once and monitored for
conformance; and
• Recognize that for the highest-consequence presentations, greater directiveness is often the
clinically safer design. An emergency disposition should be unambiguous. Risk assessment should
therefore evaluate whether the degree of directiveness is clinically appropriate to the situation,
rather than treating directiveness as inherently disfavored.

C. Patient-facing informational functions (Question 3)
We urge CDRH to preserve the balance the paper itself strikes. Patient-facing clinical information tools
meaningfully expand access: they reach people at the moment of need, including the large population
whose realistic alternative is not a clinician’s assessment but an unvetted internet search, a long nurseline queue, or no action at all. A categorical presumption that patient-facing functions are higher risk
would penalize exactly the tools most capable of closing access gaps, and would embed the medical
parentalism the paper cautions against.

Rather than a categorical shift on the consequences axis, we recommend CDRH evaluate patient-facing
functions against the presence of specific, verifiable safeguards:

• Outputs anchored to validated clinical protocols, with traceability available for independent and
institutional review;
• Deterministic escalation safety nets that trigger on emergency presentations regardless of
conversational context, and that cannot be suppressed by conversational drift;
• Structured deferral pathways that connect the user to human care (nurse line, scheduling,
emergency services) rather than terminating the interaction; and
• Communication validated for comprehension across health-literacy levels, consistent with the
paper’s element E.4.
Where these safeguards are present and demonstrated, patient-facing delivery should not, by itself,
elevate a function’s risk classification.

D. Bidirectional assessment of care escalation (Question 6)
Clearstep strongly supports assessing both under-escalation and over-escalation, and we speak to this
from sustained real-world measurement. The two error directions are not commensurable: the marginal
harm of a missed emergent presentation is categorically different from the marginal harm of an
unnecessary urgent-care visit, and the appropriate trade-off varies by chief complaint, population, and
care context.
We recommend that CDRH:

• Require a prespecified, clinically justified operating point. Manufacturers should state where
their escalation function is intentionally positioned, including any deliberate asymmetry toward
escalation for high-severity presentations, and justify that position against validated clinical
protocols and the deployment context, rather than asserting safety from one direction of error
alone.
• Require reporting in both directions, stratified by acuity. Aggregate concordance conceals
directionality. Under-escalation at high acuity is the paramount safety signal; over-escalation
should be characterized against its own harms (unnecessary utilization, patient anxiety, erosion of
trust) and evaluated for clinical justification at the acuity levels where it occurs.
• Benchmark both directions against human reference variability. Published literature and our own
evaluations show meaningful disagreement between qualified human raters on triage disposition.
Acceptable error bounds for a device should be set with reference to that documented human
baseline, not to an assumption of a single correct answer.

II. Competency-Based Premarket Evaluation (Section V; Questions 7–10, 14–16)
A. Support for the competency-based approach (Questions 7 and 8)
Clearstep supports the competency-based structure of non-clinical device benchmarking followed by
clinical confirmation, evaluated on the final user-facing device as configured for deployment.
We agree
the two-axis framework should calibrate the rigor, scope, and amount of evidence: the benchmarking
elements exercised, the difficulty and adversarial depth of test conditions, and the rung of clinical
confirmation should all be functions of grid position. We also support the paper’s recognition that
clinical confirmation need not mean a prospective clinical study in every case; the enumerated ladder of
retrospective evaluation, shadow deployment, standardized interactions, independent clinician
adjudication of real cases, and prospective study reflects methods that are practical and already in use
in responsible commercial deployments today.

B. Benchmarking elements and architectural risk controls (Question 9)
The proposed elements (S.1–S.3, E.1–E.4, R.1–R.2, A.1) are well chosen, and we particularly endorse the
treatment in R.1 of variation in safety-critical behaviors (escalation, refusal, and diagnostic conclusions)
as failures, while tolerating variation in non-safety-critical phrasing. We recommend one addition to
how the elements are applied: the framework should explicitly recognize architectural risk controls in
scoping the benchmarking burden.

GenAI-enabled devices differ fundamentally in where generative components sit. In hybrid, expertencoded architectures, safety-critical behaviors are governed by deterministic clinical logic traceable to
validated protocols; generative components handle communication and information elicitation. Such
systems can guarantee, by construction and verifiably, that semantically equivalent inputs yield identical
escalation and disposition behavior, that outputs remain within a bounded, prespecified disposition set,
and that every recommendation is traceable to a protocol pathway. A purely generative system can only
estimate those same properties statistically, across samples. Evaluation of R.1 (reproducibility), S.2
(scope maintenance), and S.3 (calibration) should permit sponsors to demonstrate these properties
through architectural verification plus targeted testing, rather than exclusively through exhaustive
sampling. This is not a lower bar; it is a different and often stronger form of evidence, and recognizing it
would encourage the industry toward designs that place generative capability where variance is
tolerable and deterministic logic where it is not.

C. Benchmark construct validity (Question 10)
We share CDRH’s concern about contamination, saturation, and limited representativeness of public
benchmarks. In our experience, the strongest evidence that a benchmark predicts real-world behavior is
a demonstrated linkage between benchmark performance and independently adjudicated real-world
performance in the same deployed configuration.
We recommend that sponsors be permitted, and
expected, to establish construct validity by:

• Building benchmark assets from clinician-authored cases graded against validated clinical
protocols, with authorship, grading rubrics, and adjudicator qualifications documented;
• Demonstrating correlation between benchmark results and real-world concordance measured
through independent clinician review of deployed interactions; and
• Complementing sponsor-developed assets with sequestered, third-party-held test sets where
available, to address optimization-to-the-test concerns.
Sponsor-developed benchmarks should not be disfavored per se: for a specific intended use, a welldocumented sponsor benchmark tied to the device’s clinical content and validated against real-world
adjudication will often have higher construct validity than a public asset built for a different purpose.
Independence concerns are better addressed through documentation, sequestration, and independent
adjudication than through exclusion.

D. Comparators and performance standards (Questions 14 and 15)
For open-ended outputs where a single correct response does not exist, we recommend:

Default to the median qualified clinician in practice, not an idealized consensus panel. Consensus
panels are valuable as a reference ceiling and for adjudicating disagreement, but requiring devices
to exceed a consensus standard that practicing clinicians themselves do not consistently meet
would hold devices to a bar unrelated to the care patients actually receive. Documented interrater variability among qualified clinicians on the same task should be measured and reported
alongside device performance.
• Permit counterfactual comparators with justification (Question 15). For many patient-facing
access functions, the realistic alternative to the device is not clinician assessment; it is self-directed
internet searching, prolonged queue times, or inaction. Where a sponsor can characterize the
counterfactual with data (for example, baseline disposition behavior of the intended population
absent the device), evaluation against that counterfactual reflects the device’s true public-health
effect. We recommend CDRH develop a methodological pathway for identifying and justifying such
comparators.
• Match the evaluation basis to the deployed workflow. Where the device operates with a clinician
in the loop, human-AI team performance is the correct basis; where it operates autonomously,
device-alone performance is. Sponsors should characterize the deployed autonomy level precisely,
and evaluation should follow it. A single device family may warrant both bases across its
deployment modes.

E. Independent third parties (Question 16)
Clearstep supports roles for qualified, independent third parties, particularly in maintaining sequestered
evaluation datasets and providing independent expert adjudication, and we speak from experience,
having voluntarily undergone independent third-party technical evaluation of our platform, including
model architecture, confidence-threshold methodology, and clinical evaluation processes. Third-party
involvement improves evidence quality and reviewer confidence. However, program design will
determine whether it broadens or narrows the market. We recommend:

Multiple accredited bodies per clinical domain, with published accreditation criteria and
transparent, capped fee schedules, so certification does not become a fixed-cost moat favoring
incumbents
;
• A sponsor right to self-assess against the same published methods, subject to FDA review, as an
alternative to third-party certification; and
• Independence criteria covering both the device sponsor and, where a third-party foundation
model is incorporated, that model’s developer, including when the adjudicator is itself an LLM; in
that case, the adjudicating model should not share a developer or lineage with the component
under evaluation.

III. Postmarket Monitoring (Section VI; Questions 18, 19, 21, 24)
A. Rebalancing premarket and postmarket evidence (Question 18)
Clearstep strongly supports accepting greater premarket uncertainty in exchange for robust postmarket
monitoring, for a reason the paper identifies: for open-ended conversational devices, no feasible
premarket test can enumerate the input space, while real-world deployment generates exactly the
evidence that premarket testing cannot.
This approach is appropriate where the sponsor commits, in
advance, to: a prespecified monitoring plan with defined sampling frames, cadences, analyses, and
thresholds; defined triggering events (including changes to any underlying model or deployment
architecture) that initiate re-benchmarking against the premarket baseline; independent clinician
adjudication of sampled real-world interactions; and defined escalation and corrective-action
procedures when thresholds are breached. It is less appropriate for fully autonomous action-taking
functions at high consequence, where premarket evidence should remain primary.

B. Feasibility of the proposed monitoring approaches (Question 19)
We can confirm from operating experience that the paper’s proposed approaches are practical at
production scale. Clearstep operates a continuous postmarket review program in which board-certified
clinicians, with cross-specialty review to reduce adjudication bias, evaluate randomly sampled realworld interactions across triage acuity tiers on a recurring (weekly to bi-weekly) cadence, alongside
automated regression evaluation of clinical content and monitoring for shifts in the input population
and interaction patterns. This program has operated for years across a deployment footprint spanning
millions of patient interactions. Our recommendations from that experience:

Sampling should be stratified by acuity and by interaction pattern, not uniform, because safety
signal concentrates in high-acuity and atypical-trajectory interactions
;
• Cadence should scale with deployment volume and change frequency, with defined triggering
events (model or content updates, new deployment contexts, detected input-population drift)
initiating off-cycle review; and
• Input-population drift deserves explicit standing as a monitored signal: changes in who reaches
and completes interactions with a device can shift its effective performance without any change to
the device itself.
Regarding machine-based supervisory agents (Question 20), we support their use as a scaling layer for
surveillance that flags candidate interactions for human review, provided the supervisory agent is itself
evaluated, its independence from the supervised components is documented, and final adjudication
authority for safety-relevant findings remains with qualified clinicians.

C. Ecosystem roles without diffusing accountability (Question 21)
Health systems, clinicians, and professional societies have essential roles in local validation, deployment
governance, and monitoring, and CDRH should encourage that participation. We ask CDRH to address
one clarification directly in any future policy: a deploying healthcare institution’s participation in
postmarket monitoring, local validation, or configuration of a device within the manufacturer’s
prespecified, validated envelope should not, by itself, expose that institution to classification as a
manufacturer or specification developer
. Absent that clarity, the institutions best positioned to
contribute monitoring evidence face a regulatory disincentive to do so, and manufacturers face pressure
to exclude them, which is the opposite of the shared-ecosystem model the paper envisions.
Manufacturer accountability is preserved by keeping the monitoring plan, thresholds, and correctiveaction obligations with the sponsor, with institutional contributions feeding that plan.

D. Changes initiated by third-party foundation model developers (Question 24)
Clearstep has direct experience with the most severe version of this risk: earlier in our history, an
upstream AI vendor on which our product depended was acquired and its service discontinued, forcing a
complete replacement of that dependency. That experience shaped an approach we believe generalizes
into the mechanisms CDRH seeks:

• Architectural containment. Keep safety-critical clinical logic outside the third-party model, so that
upstream changes affect communication quality rather than clinical dispositions. This converts an
unbounded change-control problem into a bounded one and should be recognized as a primary
mitigation in change-control review.
• Version pinning with contractual notice. Sponsors should pin model versions and secure
contractual minimum-notice windows for deprecations, with continued access to prior versions
during revalidation.
Foundation model developers do not consistently offer these terms today;
FDA’s articulation of them as expected practice would materially strengthen device manufacturers’
ability to obtain them.
• Shadow-mode gating and regression benchmarking. New model versions should run in shadow
mode against production traffic and a prespecified regression battery derived from the premarket
benchmarking baseline, with promotion to production gated on passing prespecified acceptance
criteria. Clearstep operates this pattern today.
• PCCP coverage. Foundation model version changes that pass the prespecified regression battery
within a defined envelope are a natural category for inclusion in a PCCP, enabling timely adoption
of upstream improvements without a new submission while preserving the premarket baseline as
the reference standard.

IV. Foundation Model Master Files and Agentic Systems (Section VII; Questions 25, 26)
A. Foundation Model MAFs (Question 25)
As a manufacturer that incorporates third-party foundation models in peripheral (non-safety-critical)
roles, Clearstep supports voluntary Foundation Model MAFs: structured model and system
documentation available to review teams on a reference basis would reduce duplicative diligence and
improve review consistency. Two design points are essential:

• Because participation is voluntary and developer incentives are limited, the absence of an MAF
must create no presumption against a device. Sponsor-conducted characterization of the
incorporated model on a defined battery, scoped to the model’s role in the device, should be a
fully sufficient substitute; otherwise the program would advantage sponsors with privileged
foundation-model relationships and disadvantage everyone else.
• Platform intermediaries (cloud model platforms that host and version foundation models) should
be eligible MAF holders. They are frequently the actual contractual counterparty for model access,
hold the version-control and deprecation levers, and may participate where an upstream
developer would not.

B. Agentic AI systems (Question 26)
We support the paper’s recognition that agentic systems performing care coordination, clinical
documentation, patient outreach, and workflow support may not be the focus of device oversight, and
we urge CDRH to preserve that boundary clearly
: these functions are where autonomy delivers access
and efficiency gains at low clinical risk, and ambiguity would chill beneficial deployment. For agentic
functions that are devices, we recommend acceptance criteria and oversight reflect three additions to
the A.1 element:

Mandatory human-oversight checkpoints before irreversible or high-consequence actions, verified
as non-bypassable in the deployed configuration rather than merely documented
;
Prompt-injection resistance evaluated across all input surfaces (user input, retrieved content
including EHR-derived content, and tool outputs), since agentic devices embedded in clinical
systems ingest substantial third-party content
; and
Characterized failure envelopes for tool errors and unavailability, demonstrating that the system
degrades to a safe state (deferral or human handoff) rather than improvising around a failed
dependency.

Conclusion
The discussion paper is a substantive and well-constructed foundation for regulating this category. Its
most important commitments align with what responsible manufacturers in this space already practice:
risk proportionate to function and consequence, evaluation of the deployed device as configured,
recognition that premarket testing cannot exhaust open-ended input spaces, and a meaningful role for
postmarket evidence. Clearstep would welcome the opportunity to discuss any of these comments with
CDRH, to participate in workshops or pilots developing the competency-based approach, and to share
further detail on the real-world monitoring and evaluation programs described above.
Thank you for your consideration.

Respectfully submitted,

Bilal Naved, PhD
Co-Founder & Chief Product Officer
Clearstep Inc.
bilal@clearstep.health