Richard Pescatore, DO (BellyMD)
“Holding low-risk educational software to a standard the delivery system itself does not meet protects no one.”
What they argued
Q3 patient-facing not per se elevation with envelope/deferral; Q11 confirmation tiers; Q18 support with monitoring conditions; Q24 PCCP for like-for-like model upgrades.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ11 · Clinical confirmation without a prospective trialQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ18 · Trading premarket certainty for postmarket monitoringQ24 · Third-party foundation model changesQ25 · Foundation Model Master Files
Coded positions
Consider how personalized the answer is
Consider the user and clinical context
Do not raise risk just because the user is a patient
Enforce limits on what the conversation can do
Require prospective studies for specified higher-risk uses
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Across the five cross-cutting questions
High-consequence work: Advises
The comment as filed
I am a board-certified emergency physician and founder/CEO of BellyMD, Inc., a Philadelphia digital health company building patient-facing software, including AI-supported symptom tracking and education, for disorders of gut-brain interaction (DGBI). I write as both a practicing clinician and a small manufacturer. My full comment is attached; a summary follows.
The paper’s direction is sound: the two-axis framework is workable, and the competency-based model maps naturally onto how medicine already evaluates human judgment. Key points from the attached letter:
Q1: Risk assessment should represent (a) the counterfactual care environment (the consequence of an incorrect output is measured against what would have happened without the device, which in underserved conditions like DGBI is often no professional input at all) and (b) duration of reliance (longitudinal companion tools compound error into belief; exposure time should modify risk the way dose modifies toxicity).
Q2: Judge directiveness by substance, not wording; discount boilerplate disclaimers. Auditable modifiers: specificity, personalization, imperative framing, and traceability to validated instruments and guidelines (e.g., Rome criteria, PROMIS), which should be credited as risk-reducing. Publish worked examples along the four-step gradient so small sponsors can classify functions before building.
Q3: Do not make "patient-facing" a per se risk elevation. In DGBI, the realistic alternative to a well-designed informational tool is unvetted internet content, not a clinician conversation. Credit verifiable behaviors: structurally non-directive output frames, forced deferral on treatment-change requests, escalation scaffolding tested in both directions (S.1), and health-literacy communication testing (E.4).
Q5: Assess conversational devices at the trajectory level via adversarial longitudinal simulation with standardized personas (symptom minimizers, escalation resisters). Formalize a "conversational envelope": behaviors the device must never exhibit in any trajectory, enforced and reported as violation rates. Characterize intended use by the envelope, which is testable.
Q6: As an emergency physician I see both escalation failures weekly. They are asymmetric and not commensurable; do not collapse them into one score. Sponsors should prespecify asymmetric error tolerances by clinical context and report both rates against clinician performance on the same scenarios. DGBI alarm features warrant near-zero under-escalation tolerance; bounded over-referral tolerance is acceptable and honest.
Q11: Support the confirmation ladder. For non-directive, low-consequence informational functions, retrospective evaluation plus standardized patient interactions should ordinarily suffice. Publish presumptive confirmation tiers keyed to the two-axis framework. If every conversational function requires a prospective trial, the field consolidates to incumbents and low-risk tools for unserved patients are never built.
Q14/Q15: The most consequential question in the paper. For patient-facing informational functions, a specialist panel is often the wrong comparator because it does not describe what the device displaces; in much of DGBI care, the device replaces nothing or replaces uncurated internet content. Anchor comparator selection to the realistic deployment context; recognize a "usual information environment" comparator, validated by sampling what patients encounter when they search their symptoms; reserve clinician-panel parity for functions that displace clinician judgment; where clinician comparison applies in generalist contexts, use the median generalist.
Q18: Support postmarket-weighted evidence with conditions: prespecified monitoring plans with quantitative thresholds, automated drift surveillance with sampled independent clinician adjudication, transparent reporting cadence, and demonstrated rollback capability.
Q24/Q25: Small manufacturers cannot compel foundation model developers to disclose changes. Treat version pinning as a baseline design expectation; advance Foundation Model MAFs with update-notification commitments, signaling more predictable review for devices built on MAF-holding models; endorse sponsor-side re-benchmarking gates before any model change enters production; define a PCCP category for like-for-like model version upgrades validated through prespecified re-benchmarking.
Q16: Third-party benchmarks and adjudication are appropriate, but fee structures must scale with sponsor size or certification becomes a toll gate only incumbents can pay.
Full responses with supporting reasoning are in the attached letter.
Richard Pescatore, D.O.
Founder and CEO, BellyMD, Inc.
Attachment
August 18, 2026
Submitted electronically via Regulations.gov
Re: Docket No. FDA-2026-N-7874; Considerations for the Regulation of Generative AI-Enabled
Medical Devices: Discussion Paper and Request for Feedback
To the Center for Devices and Radiological Health:
I am a board-certified emergency physician and the founder and chief executive officer of BellyMD,
Inc., a Philadelphia-based digital health company building patient-facing software, including
AI-supported symptom tracking and education, for people with disorders of gut-brain interaction
(DGBI) such as irritable bowel syndrome and functional dyspepsia. I continue to practice emergency
medicine. I write from both chairs: as a clinician who manages the downstream consequences when
patients act on bad information, and as a small manufacturer that expects to meet the standards this
paper contemplates.
The direction of the paper is sound. The two-axis framework is a workable heuristic, the
competency-based model maps naturally onto how medicine already evaluates human judgment, and
the recognition that clinical confirmation need not mean a prospective trial in every case reflects real
least-burdensome thinking. I offer responses to the discussion questions where my clinical practice and
my company's work give me direct knowledge.
Question 1: additional risk dimensions. Two dimensions deserve explicit representation. First, the
counterfactual care environment. The consequence of relying on an incorrect output is properly
measured against what would have happened without the device, and that baseline varies widely across
deployment contexts. A function deployed to a population with ready access to specialty care presents a
different risk calculus than the same function deployed to a population whose realistic alternative is no
professional input at all. Section V.D.1 and Question 15 touch this in the comparator discussion; it
belongs in the risk assessment itself. Second, duration and accumulation of reliance. A longitudinal
companion that shapes a patient's understanding of their condition over months is different from a
single-encounter tool, because error in the former compounds and becomes belief. Exposure time
should modify risk the way dose modifies toxicity.
Question 2: modifiers of directiveness. I support judging directiveness by substance and context
rather than wording, and I support discounting boilerplate disclaimers. The modifiers that can be
specified and audited are: specificity (a named drug and dose versus class-level education);
personalization (general statements versus statements tied to the user's own data); imperative framing;
and traceability, meaning whether the output is grounded in and cites validated instruments and
published guidelines. Outputs constructed from validated frameworks (in my field, the Rome
diagnostic criteria and PROMIS measures) are easier for users, sponsors, and reviewers to evaluate than
free generation, and that grounding should be credited as risk-reducing. On predictability: publish
worked examples along the four-step gradient the paper describes (general information,
contextualization, endorsement, instruction) for a handful of clinical domains. Small sponsors need to
be able to classify their own functions before they build, not after a deficiency letter.
Question 3: patient-facing functions. The paper is right that a "talk to your doctor" line does not make
an output less directive, and boilerplate should not be accepted as mitigation. But I urge CDRH not to
convert "patient-facing" into a per se elevation on the consequences axis. In DGBI, diagnostic delay is
commonly measured in years, most patients are managed without subspecialty input, and many are
managed with no professional input at all. The practical alternative to a well-designed informational
tool is not a clinician conversation; it is whatever search results and social content the patient finds
alone. The mitigations that should count are verifiable behaviors: output frames that are structurally
non-directive, teaching mechanisms and describing guideline-concordant options without endorsing
one; forced deferral whenever a user requests a treatment change; escalation scaffolding tested in both
directions (element S.1); and communication testing across health-literacy levels (element E.4). A
device that demonstrates these behaviors under adversarial testing has done something a disclaimer
never did.
Question 5: multi-turn trajectories. Trajectory-level assessment is the correct unit of analysis. A
practical method: adversarial longitudinal simulation using standardized personas, including symptom
minimizers, escalation resisters, and users who persistently solicit directive advice, scored for drift
along the directiveness continuum and for time-to-escalation across the full exchange. I suggest
formalizing a "conversational envelope": the sponsor declares the behaviors the device must never
exhibit at any point in any trajectory (for example, never proposing a medication dose change),
demonstrates enforcement under adversarial testing, and reports envelope violation rates. Intended use
for a conversational device is then characterized by its envelope, which is testable, rather than by
per-turn labels, which are not.
Question 6: escalation error in both directions. I see both failure modes in the emergency
department every week. Under-escalation injures the patient in front of the device. Over-escalation
injures the same patient through cascade testing, cost, and anxiety, and injures everyone else through
crowding and triage dilution; over time it erodes the trust that appropriate care-seeking depends on. The
two directions are real, asymmetric, and not commensurable, and they should not be collapsed into a
single utility score. Sponsors should prespecify asymmetric error tolerances justified by clinical context
and report both rates against clinician performance on the same scenario set. For DGBI functions, alarm
features (gastrointestinal bleeding, unintended weight loss, progressive dysphagia, nocturnal
symptoms) warrant near-zero tolerance for under-escalation, while a bounded, disclosed tolerance for
conservative over-referral is acceptable and honest.
Question 11: proportionate clinical confirmation. I support the confirmation ladder and the
recognition that a prospective study is not always necessary. For non-directive informational functions
in the low-consequence region of the framework, retrospective evaluation on real-world inputs plus
standardized patient interactions should ordinarily suffice, with prospective study reserved for
action-directing and action-taking functions and high-consequence domains. I ask CDRH to publish
presumptive confirmation tiers keyed to position on the two-axis framework, so a sponsor can locate its
device and know the default evidence expectation before designing a program. The economics deserve
plain statement: if every conversational function requires a prospective trial, the field consolidates to
the largest incumbents, and the low-risk informational functions that would have reached unserved
patients are never built. That outcome has a public-health cost too, and it is paid by patients who
currently receive nothing.
Questions 14 and 15: comparators. This is the most consequential question in the paper. For
informational functions deployed directly to patients, a specialist panel is often the wrong comparator
because it does not describe what the device displaces. In much of DGBI care, the device does not
replace a specialist; it replaces nothing, or it replaces uncurated internet content. I recommend: first,
comparator selection anchored to the realistic deployment context, supported by care-access data for
the intended population; second, recognition of a "usual information environment" comparator,
validated by sampling what patients in the intended population encounter when they search their
symptoms; third, reservation of clinician-panel parity for functions that displace clinician judgment,
meaning action-directing and action-taking functions or deployment inside clinical workflows; and
fourth, where clinician comparison is appropriate for generalist-shaped contexts, the median generalist
rather than a specialist panel. Holding low-risk educational software to a standard the delivery system
itself does not meet protects no one. It preserves the status quo for patients whose status quo is nothing.
Question 18: postmarket-weighted evidence. I support accepting greater premarket uncertainty in
exchange for postmarket monitoring, under conditions: a prespecified monitoring plan with quantitative
thresholds and defined triggers; automated performance and drift surveillance with sampled
adjudication by independent clinicians; a transparent reporting cadence; and demonstrated rollback
capability. A deployed conversational device generates more decision-relevant performance evidence
in a month of instrumented use than any premarket sample can contain. A regulatory posture that
credits well-built monitoring will also push manufacturers to build the instrumentation, which improves
safety independent of any submission.
Questions 24 and 25: third-party foundation models. A small manufacturer cannot compel a
foundation model developer to disclose changes or give advance notice, and contractual leverage is
concentrated in the largest sponsors. Four mechanisms would help. Treat model version pinning as a
baseline design expectation, with developers disclosing deprecation timelines. Advance the voluntary
Foundation Model MAF program with update notification commitments and healthcare-relevant
evaluation summaries as core content, and signal that devices built on MAF-holding models will see
more predictable review; that signal creates the developer's incentive to participate. Endorse
sponsor-side re-benchmarking gates, under which no underlying model change enters production until
the premarket benchmark battery has been re-run and passed, consistent with Section VI.C. And define
a PCCP category for like-for-like model version upgrades validated through that prespecified
re-benchmarking. Together these convert an uncontrollable third-party dependency into a controlled,
auditable change process inside the sponsor's quality system.
Question 16: independent third parties. Sequestered benchmark datasets and independent
adjudication panels are the right roles for qualified third parties, and ASCA and MDDT are reasonable
scaffolds. One caution: fee structures and access terms must scale with sponsor size, or third-party
certification becomes a toll gate that only incumbents can pay, with the competitive consequences the
paper itself warns against.
CDRH's instinct to evaluate these systems the way medicine evaluates clinicians, through demonstrated
competency, enforced boundaries, and supervised confirmation rather than exhaustive enumeration of
inputs, is the right one. I appreciate the agency's early engagement on this topic and would welcome the
opportunity to contribute further, including on evaluation scenarios specific to disorders of gut-brain
interaction.
Respectfully submitted,
Richard Pescatore, D.O.
Founder and Chief Executive Officer, BellyMD, Inc.
Board-Certified Emergency Physician
Philadelphia, Pennsylvania