David Wu Shi, MD (Middle Wave)
“Before we ask, “Is the AI confident?”, we need to ask, “Does the medical evidence justify that confidence?””
What they argued
Accepts FDA benchmarking (calibration, proficiency, subgroups) but adds evidence-quality calibration test using historical reversals; no position on trials, autonomy or change control.
Themes it raises
Across the five cross-cutting questions
High-consequence work: Not stated
The comment as filed
Before FDA asks whether a generative AI model is confident, it should ask whether the medical evidence justifies that confidence.
My recommendation is simple: add evidence-quality calibration to the evaluation of generative AI-enabled medical devices. A model should distinguish between precise, well-characterized evidence and large amounts of observational, surrogate, or interpretation-heavy data. More data can increase precision without fixing bias or establishing validity.
FDA could test this directly by using historical medical reversals. Give a model the literature available immediately before a major trial overturned prevailing assumptions, and evaluate whether it recognizes the limits of the evidence rather than simply reproducing the consensus.
Subgroup evaluation should likewise focus on whether a device is safe, usable, and clinically effective for each population, including people facing health-literacy, language, or technology barriers, rather than assuming that mathematical parity alone produces equitable outcomes.
The objective is not an AI that always gives an answer. It is an AI that can recognize when the evidence is weak, communicate that uncertainty clearly, and know when clinical judgment must take over.
The attached comment provides the proposed framework, historical benchmarks, examples, and supporting literature.
David W. Shi MD
Attachment
Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback
Submitted by: David Wu Shi, MD | Teach & Serve
Thank you for requesting public feedback on generative AI medical devices. FDA is considering benchmarks for
calibration, clinical proficiency, and subgroup performance.
I recommend adding one practical check: Before we ask, "Is the AI confident?", we need to ask, "Does the
medical evidence justify that confidence?"
An AI can flawlessly summarize medical literature and still deliver a dangerously incorrect answer. Why? Because
the literature itself often scales a systematic error. A trustworthy AI must recognize when it is standing on
quicksand and say so.
Dataset Volume and Validity Are Not the Same Thing
Let’s use plain terms. Precision means a measurement is consistent. A bathroom scale reading "150.2 lbs" three
times in a row is precise. Validity means the measurement reflects the actual target. If the scale is miscalibrated
and you weigh 140 lbs, the reading is beautifully precise, but reliably incorrect.
Adding more data just narrows random noise. It does nothing to make the answer valid.
Figure 1. A confidence interval only measures random error (precision), not whether the estimate is actually on target (validity).
(Source: Textbook of Epidemiology, 2nd ed., 2023).
As shown in Figure 1, a narrow confidence interval can quantify sampling precision while leaving systematic error
untouched. If a study design has important limitations, a larger dataset can make the estimate more precise
without making the underlying conclusion more valid.
Human Wishful Thinking and Citation Snowballing
Human brains love a clean story. When a biological mechanism sounds intuitive, we naturally want to believe it,
often mistaking a correlated pattern for an established cause. An AI should not inherit our cognitive soft spots; it
needs to be engineered as a dispassionate check against them – slow, methodical, and skeptical of neat
narratives.
This is why a medical AI cannot lean on an author's h-index or a paper's citation volume to justify its confidence.
Medical literature is prone to citation snowballing: once a plausible hypothesis enters the ecosystem, subsequent
papers cite it repeatedly, creating an echo chamber of apparent certainty around a fragile empirical base. The AI
must evaluate the structural integrity of the study design, not the popularity of the conclusion.
Historical Reversals
Modern medicine includes major moments where vast observational data pointed one way, only for rigorous
randomized trials to reveal the opposite. This was rarely a failure of clinician competence; it was simply the
inherent ceiling of observational data meeting our human eagerness to help patients.
● Hormone Replacement Therapy & Heart Disease (WHI): Large prospective cohorts suggested hormone
therapy halved coronary risk. The randomized Women’s Health Initiative evaluated the hypothesis and
found increased coronary heart disease and stroke. The lesson: Massive sample size cannot balance out
unmeasured confounding variables.
● Heart Rhythm Drugs (CAST): Post-heart attack premature ventricular beats correlated with sudden
death, and encainide and flecainide reliably suppressed them. Yet in the randomized CAST trial, patients
taking the drugs died at twice the rate of those on placebo. The lesson: Moving a surrogate metric in the
desired direction does not guarantee the patient benefits.
● Beta-Carotene (CARET): Diets rich in beta-carotene correlated with lower lung cancer rates. Yet when
tested as an isolated supplement in smokers, it produced a 28% increase in lung cancer and higher
mortality. The lesson: You cannot isolate an element from a healthy diet into a capsule and assume the
health effect tags along.
● Blood Sugar Optimization (ACCORD): In type 2 diabetes, higher HbA1c correlates with cardiovascular
risk. The ACCORD trial aggressively pushed HbA1c lower, but the intensive group experienced 22%
higher all-cause mortality, halting the trial. The lesson: A valid lab test becomes harmful when optimizing
the number replaces treating the patient.
Grading the Evidence
An AI should not treat all text as equal. It needs to grade the evidence:
● Characterized Ground Truth: Hardware devices use standard calibration weights. Timepieces use
an atomic clock. Medical AI needs the same thing: clinical reference datasets that act as the anchor
for reality. This is data with a short, traceable path from physical state to the electronic record.
Examples include a serum potassium level of 4.2 mmol/L measured by a calibrated electrode, or a
binary count of 30-day all-cause mortality.
● Structured probabilistic evidence: Major clinical trials that actively control for confounding variables.
● Noisy or interpretation-heavy evidence: Weak surrogate endpoints, shifting pilot research, or subjective
clinical notes.
When evidence is weak, the AI needs structural humility. It should widen its uncertainty bounds and defer to a
clinician. Crucially, a massive dataset should only increase confidence within its lane. A million rows of noisy
observational data do not magically transform into ground truth.
Reframing Fairness as Usability and Safety
This same logic applies to how we regulate subgroup performance. FDA should avoid treating subgroup
evaluation as a requirement for identical error rates across patient groups. Statistical parity and equitable clinical
outcomes are not necessarily the same thing.
A 2025 simulation of AI breast-cancer screening (Stanley et al.) found that mathematically equalizing true-positive
rates reduced the disparity partly by increasing mortality in the previously higher-performing group. We call this
"leveling down," and it conflicts directly with non-maleficence.
Worse, the mathematical fix is weak. The same study showed that simply ensuring disadvantaged groups actually
received the screening reduced disparities five times more effectively than equalizing the algorithm's odds.
FDA should frame subgroup requirements as usability and safety testing. Can patients with low health literacy,
limited English proficiency, or limited tech access safely use the device? Regulators should evaluate clinical safety
boundaries for each group independently, rather than treating mathematical parity as the primary goal when it
may lower the ceiling for some without fixing real-world barriers.
An Operational "Before-and-After" Benchmark
FDA can turn this into a practical test using historical blind spots.
Feed a model the medical literature exactly as it existed the day before the WHI or CAST trials dropped. A
passing model should not just parrot the prevailing consensus of that day. It must evaluate the structural limits of
the data and say: "This is an observation, not an established cause," or "This improves an intermediate lab
number, but hard outcome data are lacking."
Most importantly, it must hold its ground on uncertainty, even when the sheer volume of publications leans heavily
in one direction. This fits into FDA’s existing Good Machine Learning Practice framework. We are simply asking
that the quality of the evidence gets measured, not just the volume of papers.
Bottom Line
The goal is not an AI that always has a confident answer. The goal is an AI that knows the difference between
valid evidence and a mountain of noisy data – and has the discipline to warn the doctor before they both step into
the unknown.
FDA should test whether medical AI can spot the limits of our knowledge before history writes the correction.
Works Cited
1. Mac Grory B, Yeh RW, Beckman JA, et al. Observational Comparative Research in Cardiovascular and Brain Health
and Disease: A Scientific Statement From the American Heart Association. Circulation. 2026.
2. Saracci R. Epidemiology in Wonderland: Big Data and Precision Medicine. European Journal of Epidemiology. 2018.
3. Bouter L, Zeegers M, Lee T. Precision and Validity. Textbook of Epidemiology, 2nd ed. 2023.
4. Stampfer MJ, Colditz GA, Willett WC, et al. Postmenopausal Estrogen Therapy and Cardiovascular Disease. New
England Journal of Medicine. 1991.
5. LaMonte MJ, Manson JE, Anderson GL, et al. Contributions of the Women's Health Initiative to Cardiovascular
Research: JACC State-of-the-Art Review. Journal of the American College of Cardiology. 2022.
6. Col NF, Pauker SG. The Discrepancy Between Observational Studies and Randomized Trials of Menopausal
Hormone Therapy: Did Expectations Shape Experience? Annals of Internal Medicine. 2003.
7. Johansson T, Karlsson T, Bliuc D, et al. Contemporary Menopausal Hormone Therapy and Risk of Cardiovascular
Disease: Swedish Nationwide Register Based Emulated Target Trial. BMJ. 2024.
8. Echt DS, Liebson PR, Mitchell LB, et al. Mortality and Morbidity in Patients Receiving Encainide, Flecainide, or
Placebo (CAST). New England Journal of Medicine. 1991.
9. Alpha-Tocopherol, Beta Carotene Cancer Prevention Study Group. The Effect of Vitamin E and Beta Carotene on the
Incidence of Lung Cancer and Other Cancers in Male Smokers. New England Journal of Medicine. 1994.
10. Omenn GS, Goodman GE, Thornquist MD, et al. Effects of a Combination of Beta Carotene and Vitamin A on Lung
Cancer and Cardiovascular Disease (CARET). New England Journal of Medicine. 1996.
11. Abar L, Vieira AR, Aune D, et al. Blood Concentrations of Carotenoids and Retinol and Lung Cancer Risk. Cancer
Medicine. 2016.
12. Action to Control Cardiovascular Risk in Diabetes Study Group. Effects of Intensive Glucose Lowering in Type 2
Diabetes (ACCORD). New England Journal of Medicine. 2008.
13. Miller WG. The Role of Analytical Performance Specifications in International Guidelines and Standards Dealing With
Metrological Traceability in Laboratory Medicine. Clinical Chemistry and Laboratory Medicine. 2024.
14. Stanley EAM, Tsang RY, Gillett H, et al. Connecting Algorithmic Fairness and Fair Outcomes in a Sociotechnical
Simulation Case Study of AI-Assisted Healthcare. Nature Communications. 2025.
15. Chen RJ, Wang JJ, Williamson DFK, et al. Algorithmic Fairness in Artificial Intelligence for Medicine and Healthcare.
Nature Biomedical Engineering. 2023.
16. Xu Z, Li J, Yao Q, et al. Addressing Fairness Issues in Deep Learning-Based Medical Image Analysis: A Systematic
Review. npj Digital Medicine. 2024.
17. Yang Y, Zhang H, Gichoya JW, Katabi D, Ghassemi M. The Limits of Fair Medical Imaging AI in Real-World
Generalization. Nature Medicine. 2024.