← All 95 filings

David Wu Shi, MD (Middle Wave)

CliniciansClinicianFiled September 4, 20261,775 words · 1 attachmentFDA-2026-N-7874-0059
“Before we ask, “Is the AI confident?”, we need to ask, “Does the medical evidence justify that confidence?””

What they argued

RecovryAI’s one-line reading of the filing.

Accepts FDA benchmarking (calibration, proficiency, subgroups) but adds evidence-quality calibration test using historical reversals; no position on trials, autonomy or change control.

Themes it raises

2 of the 21 themes in the docket, each with the passage we counted, verbatim.
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Feed a model the medical literature exactly as it existed the day before the WHI or CAST trials dropped.”
Equity, access and under-resourced settingsFDA Q3, Q13, Q21
“Can patients with low health literacy, limited English proficiency, or limited tech access safely use the device?”

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
No position stated
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
No position stated
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
No position stated
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
No position stated
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Not stated
High-consequence work: Not stated
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Before FDA asks whether a generative AI model is confident, it should ask whether the medical evidence justifies that confidence.

My recommendation is simple: add evidence-quality calibration to the evaluation of generative AI-enabled medical devices. A model should distinguish between precise, well-characterized evidence and large amounts of observational, surrogate, or interpretation-heavy data. More data can increase precision without fixing bias or establishing validity.

FDA could test this directly by using historical medical reversals. Give a model the literature available immediately before a major trial overturned prevailing assumptions, and evaluate whether it recognizes the limits of the evidence rather than simply reproducing the consensus.

Subgroup evaluation should likewise focus on whether a device is safe, usable, and clinically effective for each population, including people facing health-literacy, language, or technology barriers, rather than assuming that mathematical parity alone produces equitable outcomes.

The objective is not an AI that always gives an answer. It is an AI that can recognize when the evidence is weak, communicate that uncertainty clearly, and know when clinical judgment must take over.

The attached comment provides the proposed framework, historical benchmarks, examples, and supporting literature.

David W. Shi MD

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback

Submitted by: David Wu Shi, MD | Teach & Serve

Thank you for requesting public feedback on generative AI medical devices. FDA is considering benchmarks for
calibration, clinical proficiency, and subgroup performance.

I recommend adding one practical check: Before we ask, "Is the AI confident?", we need to ask, "Does the
medical evidence justify that confidence?"

An AI can flawlessly summarize medical literature and still deliver a dangerously incorrect answer. Why? Because
the literature itself often scales a systematic error. A trustworthy AI must recognize when it is standing on
quicksand and say so.

Dataset Volume and Validity Are Not the Same Thing
Let’s use plain terms. Precision means a measurement is consistent. A bathroom scale reading "150.2 lbs" three
times in a row is precise. Validity means the measurement reflects the actual target. If the scale is miscalibrated
and you weigh 140 lbs, the reading is beautifully precise, but reliably incorrect.

Adding more data just narrows random noise. It does nothing to make the answer valid.

Figure 1. A confidence interval only measures random error (precision), not whether the estimate is actually on target (validity).
(Source: Textbook of Epidemiology, 2nd ed., 2023).

As shown in Figure 1, a narrow confidence interval can quantify sampling precision while leaving systematic error
untouched. If a study design has important limitations, a larger dataset can make the estimate more precise
without making the underlying conclusion more valid.
Human Wishful Thinking and Citation Snowballing
Human brains love a clean story. When a biological mechanism sounds intuitive, we naturally want to believe it,
often mistaking a correlated pattern for an established cause. An AI should not inherit our cognitive soft spots; it
needs to be engineered as a dispassionate check against them – slow, methodical, and skeptical of neat
narratives.

This is why a medical AI cannot lean on an author's h-index or a paper's citation volume to justify its confidence.
Medical literature is prone to citation snowballing: once a plausible hypothesis enters the ecosystem, subsequent
papers cite it repeatedly, creating an echo chamber of apparent certainty around a fragile empirical base. The AI
must evaluate the structural integrity of the study design, not the popularity of the conclusion.

Historical Reversals
Modern medicine includes major moments where vast observational data pointed one way, only for rigorous
randomized trials to reveal the opposite. This was rarely a failure of clinician competence; it was simply the
inherent ceiling of observational data meeting our human eagerness to help patients.

●​ Hormone Replacement Therapy & Heart Disease (WHI): Large prospective cohorts suggested hormone
therapy halved coronary risk. The randomized Women’s Health Initiative evaluated the hypothesis and
found increased coronary heart disease and stroke. The lesson: Massive sample size cannot balance out
unmeasured confounding variables.
●​ Heart Rhythm Drugs (CAST): Post-heart attack premature ventricular beats correlated with sudden
death, and encainide and flecainide reliably suppressed them. Yet in the randomized CAST trial, patients
taking the drugs died at twice the rate of those on placebo. The lesson: Moving a surrogate metric in the
desired direction does not guarantee the patient benefits.
●​ Beta-Carotene (CARET): Diets rich in beta-carotene correlated with lower lung cancer rates. Yet when
tested as an isolated supplement in smokers, it produced a 28% increase in lung cancer and higher
mortality. The lesson: You cannot isolate an element from a healthy diet into a capsule and assume the
health effect tags along.
●​ Blood Sugar Optimization (ACCORD): In type 2 diabetes, higher HbA1c correlates with cardiovascular
risk. The ACCORD trial aggressively pushed HbA1c lower, but the intensive group experienced 22%
higher all-cause mortality, halting the trial. The lesson: A valid lab test becomes harmful when optimizing
the number replaces treating the patient.

Grading the Evidence
An AI should not treat all text as equal. It needs to grade the evidence:

●​ Characterized Ground Truth: Hardware devices use standard calibration weights. Timepieces use
an atomic clock. Medical AI needs the same thing: clinical reference datasets that act as the anchor
for reality. This is data with a short, traceable path from physical state to the electronic record.
Examples include a serum potassium level of 4.2 mmol/L measured by a calibrated electrode, or a
binary count of 30-day all-cause mortality.

●​ Structured probabilistic evidence: Major clinical trials that actively control for confounding variables.
●​ Noisy or interpretation-heavy evidence: Weak surrogate endpoints, shifting pilot research, or subjective
clinical notes.

When evidence is weak, the AI needs structural humility. It should widen its uncertainty bounds and defer to a
clinician. Crucially, a massive dataset should only increase confidence within its lane. A million rows of noisy
observational data do not magically transform into ground truth.

Reframing Fairness as Usability and Safety
This same logic applies to how we regulate subgroup performance. FDA should avoid treating subgroup
evaluation as a requirement for identical error rates across patient groups. Statistical parity and equitable clinical
outcomes are not necessarily the same thing.

A 2025 simulation of AI breast-cancer screening (Stanley et al.) found that mathematically equalizing true-positive
rates reduced the disparity partly by increasing mortality in the previously higher-performing group. We call this
"leveling down," and it conflicts directly with non-maleficence.

Worse, the mathematical fix is weak. The same study showed that simply ensuring disadvantaged groups actually
received the screening reduced disparities five times more effectively than equalizing the algorithm's odds.

FDA should frame subgroup requirements as usability and safety testing. Can patients with low health literacy,
limited English proficiency, or limited tech access safely use the device?
Regulators should evaluate clinical safety
boundaries for each group independently, rather than treating mathematical parity as the primary goal when it
may lower the ceiling for some without fixing real-world barriers.

An Operational "Before-and-After" Benchmark
FDA can turn this into a practical test using historical blind spots.

Feed a model the medical literature exactly as it existed the day before the WHI or CAST trials dropped. A
passing model should not just parrot the prevailing consensus of that day. It must evaluate the structural limits of
the data and say: "This is an observation, not an established cause," or "This improves an intermediate lab
number, but hard outcome data are lacking."

Most importantly, it must hold its ground on uncertainty, even when the sheer volume of publications leans heavily
in one direction. This fits into FDA’s existing Good Machine Learning Practice framework. We are simply asking
that the quality of the evidence gets measured, not just the volume of papers.

Bottom Line
The goal is not an AI that always has a confident answer. The goal is an AI that knows the difference between
valid evidence and a mountain of noisy data – and has the discipline to warn the doctor before they both step into
the unknown.

FDA should test whether medical AI can spot the limits of our knowledge before history writes the correction.
Works Cited
1.​ Mac Grory B, Yeh RW, Beckman JA, et al. Observational Comparative Research in Cardiovascular and Brain Health
and Disease: A Scientific Statement From the American Heart Association. Circulation. 2026.
2.​ Saracci R. Epidemiology in Wonderland: Big Data and Precision Medicine. European Journal of Epidemiology. 2018.
3.​ Bouter L, Zeegers M, Lee T. Precision and Validity. Textbook of Epidemiology, 2nd ed. 2023.
4.​ Stampfer MJ, Colditz GA, Willett WC, et al. Postmenopausal Estrogen Therapy and Cardiovascular Disease. New
England Journal of Medicine. 1991.
5.​ LaMonte MJ, Manson JE, Anderson GL, et al. Contributions of the Women's Health Initiative to Cardiovascular
Research: JACC State-of-the-Art Review. Journal of the American College of Cardiology. 2022.
6.​ Col NF, Pauker SG. The Discrepancy Between Observational Studies and Randomized Trials of Menopausal
Hormone Therapy: Did Expectations Shape Experience? Annals of Internal Medicine. 2003.
7.​ Johansson T, Karlsson T, Bliuc D, et al. Contemporary Menopausal Hormone Therapy and Risk of Cardiovascular
Disease: Swedish Nationwide Register Based Emulated Target Trial. BMJ. 2024.
8.​ Echt DS, Liebson PR, Mitchell LB, et al. Mortality and Morbidity in Patients Receiving Encainide, Flecainide, or
Placebo (CAST). New England Journal of Medicine. 1991.
9.​ Alpha-Tocopherol, Beta Carotene Cancer Prevention Study Group. The Effect of Vitamin E and Beta Carotene on the
Incidence of Lung Cancer and Other Cancers in Male Smokers. New England Journal of Medicine. 1994.
10.​ Omenn GS, Goodman GE, Thornquist MD, et al. Effects of a Combination of Beta Carotene and Vitamin A on Lung
Cancer and Cardiovascular Disease (CARET). New England Journal of Medicine. 1996.
11.​ Abar L, Vieira AR, Aune D, et al. Blood Concentrations of Carotenoids and Retinol and Lung Cancer Risk. Cancer
Medicine. 2016.
12.​ Action to Control Cardiovascular Risk in Diabetes Study Group. Effects of Intensive Glucose Lowering in Type 2
Diabetes (ACCORD). New England Journal of Medicine. 2008.
13.​ Miller WG. The Role of Analytical Performance Specifications in International Guidelines and Standards Dealing With
Metrological Traceability in Laboratory Medicine. Clinical Chemistry and Laboratory Medicine. 2024.
14.​ Stanley EAM, Tsang RY, Gillett H, et al. Connecting Algorithmic Fairness and Fair Outcomes in a Sociotechnical
Simulation Case Study of AI-Assisted Healthcare. Nature Communications. 2025.
15.​ Chen RJ, Wang JJ, Williamson DFK, et al. Algorithmic Fairness in Artificial Intelligence for Medicine and Healthcare.
Nature Biomedical Engineering. 2023.
16.​ Xu Z, Li J, Yao Q, et al. Addressing Fairness Issues in Deep Learning-Based Medical Image Analysis: A Systematic
Review. npj Digital Medicine. 2024.
17.​ Yang Y, Zhang H, Gichoya JW, Katabi D, Ghassemi M. The Limits of Fair Medical Imaging AI in Real-World
Generalization. Nature Medicine. 2024.