← All 95 filings

Richard Pescatore, DO (BellyMD)

IndustryStartupFiled August 18, 20262,146 words · 1 attachmentFDA-2026-N-7874-0006
“Holding low-risk educational software to a standard the delivery system itself does not meet protects no one.”

What they argued

RecovryAI’s one-line reading of the filing.

Q3 patient-facing not per se elevation with envelope/deferral; Q11 confirmation tiers; Q18 support with monitoring conditions; Q24 PCCP for like-for-like model upgrades.

Themes it raises

11 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Exposure time should modify risk the way dose modifies toxicity.”
Whether the user can judge the outputFDA Q3, Q4
“But I urge CDRH not to convert "patient-facing" into a per se elevation on the consequences axis.”
Escalating too little and too muchFDA Q6
“The two directions are real, asymmetric, and not commensurable, and they should not be collapsed into a single utility score.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“CDRH's instinct to evaluate these systems the way medicine evaluates clinicians, through demonstrated competency, enforced boundaries, and supervised confirmation rather than exhaustive enumeration of inputs, is the right one.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Sequestered benchmark datasets and independent adjudication panels are the right roles for qualified third parties, and ASCA and MDDT are reasonable scaffolds.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“For non-directive informational functions in the low-consequence region of the framework, retrospective evaluation on real-world inputs plus standardized patient interactions should ordinarily suffice, with prospective study reserved for action-directing and action-taking functions and high-consequence domains.”
Trading premarket certainty for postmarket monitoringFDA Q18
“I support accepting greater premarket uncertainty in exchange for postmarket monitoring, under conditions: a prespecified monitoring plan with quantitative thresholds and defined triggers; automated performance and drift surveillance with sampled adjudication by independent clinicians; a transparent reporting cadence; and demonstrated rollback capability.”
Watching the device after it shipsFDA Q19, Q20
“A deployed conversational device generates more decision-relevant performance evidence in a month of instrumented use than any premarket sample can contain.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“And define a PCCP category for like-for-like model version upgrades validated through that prespecified re-benchmarking.”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“A function deployed to a population with ready access to specialty care presents a different risk calculus than the same function deployed to a population whose realistic alternative is no professional input at all.”
What the rules cost sponsors and the marketNot asked by the FDA
“The economics deserve plain statement: if every conversational function requires a prospective trial, the field consolidates to the largest incumbents, and the low-risk informational functions that would have reached unserved patients are never built.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ11 · Clinical confirmation without a prospective trialQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ18 · Trading premarket certainty for postmarket monitoringQ24 · Third-party foundation model changesQ25 · Foundation Model Master Files

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
Keep it, but add or change elements
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider the wording and specificity
Consider how personalized the answer is
Consider the user and clinical context
Q3When clinical information goes straight to the patient, does the risk change, and what safeguards help without underestimating patients?
Base risk on the task and available safeguards
Do not raise risk just because the user is a patient
Q5How is risk assessed when a conversation starts with non-directive information and drifts into action-directing?
Test whole conversations, not isolated answers
Enforce limits on what the conversation can do
Q6How should under-escalation be weighed against over-escalation?
Set the trade-off for the clinical context
Q11When can a device be confirmed without a prospective clinical study, and what earns that lighter path?
Some uses can be confirmed without a prospective study
Require prospective studies for specified higher-risk uses
Q14For open-ended AI outputs, who is the performance comparator: a clinician panel, generalists, specialists, or the human-AI team?
Compare with what happens without the device
Q15Could the AI be measured against what would have happened without it: unaided judgment, a delayed specialist, or no intervention?
Use that comparator, with conditions
Q16What role should independent third parties play?
Use independent parties to hold or maintain test assets
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Q18Can greater premarket uncertainty about a GenAI device’s benefit-risk profile be accepted through greater reliance on postmarket monitoring?
Allow it only under defined conditions
Q24When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?
Identify and control the model version in use
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Q25Would voluntary Foundation Model Master Files be practical, and useful in premarket review?
Use them if specified conditions are met

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
Supports with conditions
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Directs
High-consequence work: Advises
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

I am a board-certified emergency physician and founder/CEO of BellyMD, Inc., a Philadelphia digital health company building patient-facing software, including AI-supported symptom tracking and education, for disorders of gut-brain interaction (DGBI). I write as both a practicing clinician and a small manufacturer. My full comment is attached; a summary follows.

The paper’s direction is sound: the two-axis framework is workable, and the competency-based model maps naturally onto how medicine already evaluates human judgment. Key points from the attached letter:

Q1: Risk assessment should represent (a) the counterfactual care environment (the consequence of an incorrect output is measured against what would have happened without the device, which in underserved conditions like DGBI is often no professional input at all) and (b) duration of reliance (longitudinal companion tools compound error into belief; exposure time should modify risk the way dose modifies toxicity).

Q2: Judge directiveness by substance, not wording; discount boilerplate disclaimers. Auditable modifiers: specificity, personalization, imperative framing, and traceability to validated instruments and guidelines (e.g., Rome criteria, PROMIS), which should be credited as risk-reducing. Publish worked examples along the four-step gradient so small sponsors can classify functions before building.

Q3: Do not make "patient-facing" a per se risk elevation. In DGBI, the realistic alternative to a well-designed informational tool is unvetted internet content, not a clinician conversation. Credit verifiable behaviors: structurally non-directive output frames, forced deferral on treatment-change requests, escalation scaffolding tested in both directions (S.1), and health-literacy communication testing (E.4).

Q5: Assess conversational devices at the trajectory level via adversarial longitudinal simulation with standardized personas (symptom minimizers, escalation resisters). Formalize a "conversational envelope": behaviors the device must never exhibit in any trajectory, enforced and reported as violation rates. Characterize intended use by the envelope, which is testable.

Q6: As an emergency physician I see both escalation failures weekly. They are asymmetric and not commensurable; do not collapse them into one score. Sponsors should prespecify asymmetric error tolerances by clinical context and report both rates against clinician performance on the same scenarios. DGBI alarm features warrant near-zero under-escalation tolerance; bounded over-referral tolerance is acceptable and honest.

Q11: Support the confirmation ladder. For non-directive, low-consequence informational functions, retrospective evaluation plus standardized patient interactions should ordinarily suffice. Publish presumptive confirmation tiers keyed to the two-axis framework. If every conversational function requires a prospective trial, the field consolidates to incumbents and low-risk tools for unserved patients are never built.

Q14/Q15: The most consequential question in the paper. For patient-facing informational functions, a specialist panel is often the wrong comparator because it does not describe what the device displaces; in much of DGBI care, the device replaces nothing or replaces uncurated internet content. Anchor comparator selection to the realistic deployment context; recognize a "usual information environment" comparator, validated by sampling what patients encounter when they search their symptoms; reserve clinician-panel parity for functions that displace clinician judgment; where clinician comparison applies in generalist contexts, use the median generalist.

Q18: Support postmarket-weighted evidence with conditions: prespecified monitoring plans with quantitative thresholds, automated drift surveillance with sampled independent clinician adjudication, transparent reporting cadence, and demonstrated rollback capability.

Q24/Q25: Small manufacturers cannot compel foundation model developers to disclose changes. Treat version pinning as a baseline design expectation; advance Foundation Model MAFs with update-notification commitments, signaling more predictable review for devices built on MAF-holding models; endorse sponsor-side re-benchmarking gates before any model change enters production; define a PCCP category for like-for-like model version upgrades validated through prespecified re-benchmarking.

Q16: Third-party benchmarks and adjudication are appropriate, but fee structures must scale with sponsor size or certification becomes a toll gate only incumbents can pay.

Full responses with supporting reasoning are in the attached letter.

Richard Pescatore, D.O.
Founder and CEO, BellyMD, Inc.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

August 18, 2026

Submitted electronically via Regulations.gov

Re: Docket No. FDA-2026-N-7874; Considerations for the Regulation of Generative AI-Enabled
Medical Devices: Discussion Paper and Request for Feedback

To the Center for Devices and Radiological Health:

I am a board-certified emergency physician and the founder and chief executive officer of BellyMD,
Inc., a Philadelphia-based digital health company building patient-facing software, including
AI-supported symptom tracking and education, for people with disorders of gut-brain interaction
(DGBI) such as irritable bowel syndrome and functional dyspepsia. I continue to practice emergency
medicine. I write from both chairs: as a clinician who manages the downstream consequences when
patients act on bad information, and as a small manufacturer that expects to meet the standards this
paper contemplates.

The direction of the paper is sound. The two-axis framework is a workable heuristic, the
competency-based model maps naturally onto how medicine already evaluates human judgment, and
the recognition that clinical confirmation need not mean a prospective trial in every case reflects real
least-burdensome thinking. I offer responses to the discussion questions where my clinical practice and
my company's work give me direct knowledge.

Question 1: additional risk dimensions. Two dimensions deserve explicit representation. First, the
counterfactual care environment. The consequence of relying on an incorrect output is properly
measured against what would have happened without the device, and that baseline varies widely across
deployment contexts. A function deployed to a population with ready access to specialty care presents a
different risk calculus than the same function deployed to a population whose realistic alternative is no
professional input at all.
Section V.D.1 and Question 15 touch this in the comparator discussion; it
belongs in the risk assessment itself. Second, duration and accumulation of reliance. A longitudinal
companion that shapes a patient's understanding of their condition over months is different from a
single-encounter tool, because error in the former compounds and becomes belief. Exposure time
should modify risk the way dose modifies toxicity.

Question 2: modifiers of directiveness. I support judging directiveness by substance and context
rather than wording, and I support discounting boilerplate disclaimers. The modifiers that can be
specified and audited are: specificity (a named drug and dose versus class-level education);
personalization (general statements versus statements tied to the user's own data); imperative framing;
and traceability, meaning whether the output is grounded in and cites validated instruments and
published guidelines. Outputs constructed from validated frameworks (in my field, the Rome
diagnostic criteria and PROMIS measures) are easier for users, sponsors, and reviewers to evaluate than
free generation, and that grounding should be credited as risk-reducing. On predictability: publish
worked examples along the four-step gradient the paper describes (general information,
contextualization, endorsement, instruction) for a handful of clinical domains. Small sponsors need to
be able to classify their own functions before they build, not after a deficiency letter.

Question 3: patient-facing functions. The paper is right that a "talk to your doctor" line does not make
an output less directive, and boilerplate should not be accepted as mitigation. But I urge CDRH not to
convert "patient-facing" into a per se elevation on the consequences axis.
In DGBI, diagnostic delay is
commonly measured in years, most patients are managed without subspecialty input, and many are
managed with no professional input at all. The practical alternative to a well-designed informational
tool is not a clinician conversation; it is whatever search results and social content the patient finds
alone. The mitigations that should count are verifiable behaviors: output frames that are structurally
non-directive, teaching mechanisms and describing guideline-concordant options without endorsing
one; forced deferral whenever a user requests a treatment change; escalation scaffolding tested in both
directions (element S.1); and communication testing across health-literacy levels (element E.4). A
device that demonstrates these behaviors under adversarial testing has done something a disclaimer
never did.

Question 5: multi-turn trajectories. Trajectory-level assessment is the correct unit of analysis. A
practical method: adversarial longitudinal simulation using standardized personas, including symptom
minimizers, escalation resisters, and users who persistently solicit directive advice, scored for drift
along the directiveness continuum and for time-to-escalation across the full exchange. I suggest
formalizing a "conversational envelope": the sponsor declares the behaviors the device must never
exhibit at any point in any trajectory (for example, never proposing a medication dose change),
demonstrates enforcement under adversarial testing, and reports envelope violation rates. Intended use
for a conversational device is then characterized by its envelope, which is testable, rather than by
per-turn labels, which are not.

Question 6: escalation error in both directions. I see both failure modes in the emergency
department every week. Under-escalation injures the patient in front of the device. Over-escalation
injures the same patient through cascade testing, cost, and anxiety, and injures everyone else through
crowding and triage dilution; over time it erodes the trust that appropriate care-seeking depends on. The
two directions are real, asymmetric, and not commensurable, and they should not be collapsed into a
single utility score.
Sponsors should prespecify asymmetric error tolerances justified by clinical context
and report both rates against clinician performance on the same scenario set. For DGBI functions, alarm
features (gastrointestinal bleeding, unintended weight loss, progressive dysphagia, nocturnal
symptoms) warrant near-zero tolerance for under-escalation, while a bounded, disclosed tolerance for
conservative over-referral is acceptable and honest.

Question 11: proportionate clinical confirmation. I support the confirmation ladder and the
recognition that a prospective study is not always necessary. For non-directive informational functions
in the low-consequence region of the framework, retrospective evaluation on real-world inputs plus
standardized patient interactions should ordinarily suffice, with prospective study reserved for
action-directing and action-taking functions and high-consequence domains.
I ask CDRH to publish
presumptive confirmation tiers keyed to position on the two-axis framework, so a sponsor can locate its
device and know the default evidence expectation before designing a program. The economics deserve
plain statement: if every conversational function requires a prospective trial, the field consolidates to
the largest incumbents, and the low-risk informational functions that would have reached unserved
patients are never built.
That outcome has a public-health cost too, and it is paid by patients who
currently receive nothing.

Questions 14 and 15: comparators. This is the most consequential question in the paper. For
informational functions deployed directly to patients, a specialist panel is often the wrong comparator
because it does not describe what the device displaces. In much of DGBI care, the device does not
replace a specialist; it replaces nothing, or it replaces uncurated internet content. I recommend: first,
comparator selection anchored to the realistic deployment context, supported by care-access data for
the intended population; second, recognition of a "usual information environment" comparator,
validated by sampling what patients in the intended population encounter when they search their
symptoms; third, reservation of clinician-panel parity for functions that displace clinician judgment,
meaning action-directing and action-taking functions or deployment inside clinical workflows; and
fourth, where clinician comparison is appropriate for generalist-shaped contexts, the median generalist
rather than a specialist panel. Holding low-risk educational software to a standard the delivery system
itself does not meet protects no one. It preserves the status quo for patients whose status quo is nothing.

Question 18: postmarket-weighted evidence. I support accepting greater premarket uncertainty in
exchange for postmarket monitoring, under conditions: a prespecified monitoring plan with quantitative
thresholds and defined triggers; automated performance and drift surveillance with sampled
adjudication by independent clinicians; a transparent reporting cadence; and demonstrated rollback
capability.
A deployed conversational device generates more decision-relevant performance evidence
in a month of instrumented use than any premarket sample can contain.
A regulatory posture that
credits well-built monitoring will also push manufacturers to build the instrumentation, which improves
safety independent of any submission.

Questions 24 and 25: third-party foundation models. A small manufacturer cannot compel a
foundation model developer to disclose changes or give advance notice, and contractual leverage is
concentrated in the largest sponsors. Four mechanisms would help. Treat model version pinning as a
baseline design expectation, with developers disclosing deprecation timelines. Advance the voluntary
Foundation Model MAF program with update notification commitments and healthcare-relevant
evaluation summaries as core content, and signal that devices built on MAF-holding models will see
more predictable review; that signal creates the developer's incentive to participate. Endorse
sponsor-side re-benchmarking gates, under which no underlying model change enters production until
the premarket benchmark battery has been re-run and passed, consistent with Section VI.C. And define
a PCCP category for like-for-like model version upgrades validated through that prespecified
re-benchmarking.
Together these convert an uncontrollable third-party dependency into a controlled,
auditable change process inside the sponsor's quality system.
Question 16: independent third parties. Sequestered benchmark datasets and independent
adjudication panels are the right roles for qualified third parties, and ASCA and MDDT are reasonable
scaffolds.
One caution: fee structures and access terms must scale with sponsor size, or third-party
certification becomes a toll gate that only incumbents can pay, with the competitive consequences the
paper itself warns against.

CDRH's instinct to evaluate these systems the way medicine evaluates clinicians, through demonstrated
competency, enforced boundaries, and supervised confirmation rather than exhaustive enumeration of
inputs, is the right one.
I appreciate the agency's early engagement on this topic and would welcome the
opportunity to contribute further, including on evaluation scenarios specific to disorders of gut-brain
interaction.

Respectfully submitted,

Richard Pescatore, D.O.
Founder and Chief Executive Officer, BellyMD, Inc.
Board-Certified Emergency Physician
Philadelphia, Pennsylvania