Total Medical Imaging, LLC and RADIN, LLC (Alejandro Bugnone)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
Total Medical Imaging, LLC and RADIN, LLC jointly submit the attached comment on Docket No. FDA-2026-N-7874. Please see the attached comment for the commenters’ complete analysis.
Attachment
TOTAL MEDICAL IMAGING LLC RADIN LLC
FOOD AND DRUG ADMINISTRATION
Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request
for Feedback
Joint Comment Regarding Risk-Proportionate Regulation
of Radiologist-Supervised Vision-Language Model
Report Pre-Drafting
Submitted jointly by Total Medical Imaging, LLC and RADIN, LLC
Submitted by Total Medical Imaging, LLC and RADIN, LLC
Representative Alejandro Bugnone, M.D.
Chief Executive Officer, Total Medical Imaging, LLC;
Roles Chief Executive Officer, RADIN, LLC; Practicing
Radiologist
Submission date September 25, 2026
Submission method Electronic submission to the applicable FDA docket
Joint clinical-practice and technology-developer perspective
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 1
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
I. Executive Summary
Total Medical Imaging, LLC ("TMI") and RADIN, LLC ("RADIN") respectfully submit this joint comment
in response to the Food and Drug Administration ("FDA" or "Agency") discussion paper concerning
generative artificial intelligence ("GenAI")-enabled medical devices. The discussion paper appropriately
seeks a risk-proportionate, least-burdensome framework and identifies supervision, reversibility, downstream
safeguards, specialist expertise, traceability to primary source information, human-AI team performance, and
postmarket monitoring as potentially important regulatory considerations.1
TMI and RADIN support appropriate FDA oversight when an artificial intelligence system autonomously
diagnoses a patient, communicates an unverified diagnostic conclusion to a clinician or patient, initiates or
directs clinical action, or otherwise places an AI-generated interpretation into the patient-care pathway before
meaningful specialist review.
The Agency should distinguish those uses from a materially different and substantially lower-risk function: a
vision-language model ("VLM") that generates only private, non-final textual work product for the specialist
radiologist who is independently interpreting the underlying medical images. In that workflow, the radiologist
has the complete primary source material, possesses the relevant specialty expertise, may edit or reject every
part of the generated text, and exclusively controls whether a report is issued. No unverified AI output is
available to a clinician, patient, or other downstream decision-maker.
U.S. Food & Drug Admin., Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for
Feedback 1, 5-9, 18-21 (Aug. 18, 2026), https://www.fda.gov/media/194242/download [hereinafter FDA GenAI Discussion Paper]. The paper
is for discussion only and does not state draft or final FDA policy.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 2
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Because section 520(o)(1)(E) of the Federal Food, Drug, and Cosmetic Act ("FD&C Act") generally does not
exclude software that analyzes medical images from the device definition, TMI and RADIN do not ask FDA
to declare every supervised VLM pre-drafting function categorically outside Agency jurisdiction.2 Instead, we
request that FDA establish a narrowly bounded policy of premarket enforcement discretion for a defined
category of Radiologist-Supervised Generative Pre-Drafting ("RSG-PD") functions. Section 520(o)(4)(A)
expressly preserves FDA authority to exercise enforcement discretion with respect to device software
functions.3 FDA has used function-specific enforcement-discretion policies where a software function meets
the device definition but presents lower patient risk.4
Requested policy in one sentence
FDA should exercise enforcement discretion with respect to premarket authorization requirements for a
generative artificial intelligence software function that analyzes medical images solely to generate
private, non-final report work product for review by the qualified specialist physician independently
interpreting those images, when enforceable technical controls, supported by procedural controls, prevent
the output from being finalized, communicated, acted upon, or otherwise entering the patient-care
pathway before affirmative specialist review.
The corresponding higher-risk pathway should remain subject to appropriate premarket oversight: VLM
output that is communicated to a treating clinician or patient, represented as a preliminary or final
interpretation, used to direct care, or otherwise relied upon before specialist radiologist verification.
This distinction does not ignore automation bias, hallucination, drift, subgroup performance, cybersecurity, or
other legitimate risks. It addresses those risks through human-factors evaluation, local validation,
sequestration of unverified output, ongoing monitoring, auditability, and controlled version management. The
relevant performance unit should be the final human-AI system in its intended workflow, not an isolated
foundation model tested as though it were intended to operate autonomously. Evaluation of that human-AI
system should measure both directions of effect: errors that assistance introduces and errors that a radiologist,
prompted by a structured pre-draft, detects and corrects that might otherwise have been omitted.
FDA's two-axis framework should expressly name four auditable modifiers: output disposition, clinical
persistence and reversibility, independent verifiability against primary source information, and whether
specialist review is architecturally required or merely expected. These variables determine whether an
erroneous output can independently reach a downstream decision-maker. The competency-based framework
should likewise recognize supervised operation as a distinct evidentiary stage between non-clinical testing
and independent or care-team-facing use.
II. Statement of Interest and Transparency
TMI is a Florida-based radiology practice whose more than 60 board-certified radiologists provide diagnostic
interpretation nationwide across hospital, emergency department, and outpatient imaging settings, at an
annual volume on the order of one million examinations. RADIN develops and commercially distributes a
radiology operating platform providing image management, radiology information system functions, dictation
and reporting, worklist orchestration, credentialing-aware study assignment, non-diagnostic workflow
Federal Food, Drug, and Cosmetic Act section 520(o)(1)(E), 21 U.S.C. section 360j(o)(1)(E); U.S. Food & Drug Admin., Clinical Decision
Support Software: Guidance for Industry and Food and Drug Administration Staff 5-7, 21-22 (Jan. 29, 2026),
https://www.fda.gov/media/109618/download [hereinafter FDA CDS Guidance].
FD&C Act section 520(o)(4)(A), 21 U.S.C. section 360j(o)(4)(A) (preserving FDA authority to exercise enforcement discretion as to a
device software function).
U.S. Food & Drug Admin., Policy for Device Software Functions and Mobile Medical Applications 2, 13-14, 24-25 (Sept. 2022),
https://www.fda.gov/media/80958/download; FDA CDS Guidance, supra note 2, at 5.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 3
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
automation, quality-assurance tooling, and follow-up tracking. The organizations are affiliated under common
executive leadership.
Alejandro Bugnone, M.D., is a board-certified practicing radiologist and serves as chief executive of both
organizations.
RADIN has begun research and development involving VLMs and fine-tuning using TMI data under
applicable privacy, contractual, security, and data-governance controls. No VLM-based image-to-report predrafting function arising from that work is currently deployed in clinical practice, and the commenters make
no representation that any particular model is ready for clinical use.
The commenters have a direct prospective commercial and operational interest in the regulatory framework.
We disclose that interest expressly. Our position is not that a VLM becomes safe merely because a physician
appears somewhere in the workflow. Rather, regulation should be proportionate to the actual intended use,
actual user, primary clinical source material, actual route by which an error could affect a patient, and actual
safeguards in the final configured deployment.
TMI and RADIN have separately submitted a comment in Docket No. FDA-2026-P-9175 opposing in part
the Mosaic Clinical Technologies citizen petition. This comment addresses the affirmative framework
questions presented in Docket No. FDA-2026-N-7874.
III. Terminology: A Private Pre-Draft Is Not a Preliminary Clinical Report
For purposes of this comment, an RSG-PD function is a software function that analyzes medical images
solely to generate private, non-final textual work product for use by the same qualified radiologist who
independently interprets the source images and has exclusive authority to issue the clinical report. FDA
should define this category by auditable deployment conditions rather than by model architecture or the
amount of downstream adaptation. Each condition should be verifiable from system configuration, access
controls, and audit logs.
A qualifying pre-draft:
• is available only within the interpreting radiologist's controlled workspace;
• is conspicuously identified as AI-generated, unverified work product;
• does not become part of the final medical record unless and until the radiologist affirmatively issues a report;
• is not accessible to a referring clinician, patient, or other downstream decision-maker;
• cannot be automatically signed, finalized, communicated, or relied upon;
• cannot independently trigger treatment, triage, critical-result communication, order entry, patient disposition,
or any other clinical action; and
• may be edited, deleted, completely rejected, or replaced by a report dictated from a blank state.
A preliminary report is different. If an AI-generated interpretation is transmitted or made available to an
emergency physician, referring clinician, patient, or other user who may rely upon it before specialist
verification, it is a clinical communication regardless of whether the interface labels it "draft," "preliminary,"
"suggested," or "informational." Regulatory analysis should follow actual access, reliance, and workflow, not
nomenclature.
IV. Legal and Policy Basis for a Risk-Proportionate Enforcement-Discretion Category
A. Image-analysis functions may remain device functions
Section 520(o)(1)(E) of the FD&C Act excludes certain clinical decision support functions from the statutory
device definition only when four criteria are satisfied. One criterion requires that the function not be intended
to acquire, process, or analyze a medical image, an in vitro diagnostic signal, or a pattern or signal from a
signal-acquisition system. FDA's January 2026 Clinical Decision Support Software guidance follows that
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 4
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
statutory structure and treats image-analysis functions as device software functions even when their output is
presented to a healthcare professional.5
The narrow and legally sustainable request is therefore not a blanket non-device declaration. It is a policy of
premarket enforcement discretion for a low-risk final function whose unverified output is prevented from
entering the patient-care pathway before meaningful specialist verification.
Existing FDA classifications recognize physician-assistive image-analysis functions and may require both
standalone characterization and assisted-reader performance evidence. This comment does not assume that
current oversight uniformly treats supervised software as autonomous. It asks FDA to consider whether the
additional containment, supervision, validation, and monitoring conditions proposed here warrant a distinct
enforcement-discretion policy or streamlined pathway.6
B. FDA retains enforcement-discretion authority
Congress expressly preserved FDA's authority to exercise enforcement discretion with respect to device
software functions.7 FDA has repeatedly used this authority in its digital-health policies where a function’s
intended use, deployment safeguards, and risk to patients do not warrant active enforcement of premarket
authorization requirements.8 Accordingly, FDA need not determine that radiologist-supervised VLM predrafting is categorically outside the device definition in order to establish a narrowly defined enforcementdiscretion policy for this lower-risk workflow.
A narrowly defined RSG-PD category would retain FDA jurisdiction, allow the Agency to specify limiting
conditions, and preserve enforcement authority where a sponsor markets beyond those conditions or where
evidence shows unacceptable risk. It would avoid the false choice between declaring all image-analyzing
VLMs unregulated and imposing an autonomous-diagnostic evidentiary burden on every private specialist
drafting aid.
C. The Practice-of-Medicine Provision Supports Physician Discretion but Does Not Independently
Resolve the Regulatory Status of the Commercial Software Function
Section 1006 of the FD&C Act protects a healthcare practitioner’s authority to prescribe or administer a
legally marketed device to a patient within a legitimate practitioner-patient relationship. It does not, standing
alone, authorize a manufacturer to commercially distribute a device that otherwise requires FDA
authorization. TMI and RADIN therefore do not contend that radiologist supervision categorically removes a
commercially distributed VLM function from FDA jurisdiction. The practice-of-medicine provision remains
relevant, however, because it reflects Congress’s intent to preserve qualified physicians’ authority to exercise
independent clinical judgment in caring for their patients. In the proposed pre-drafting workflow, the
radiologist independently interprets the source images, determines whether the AI-generated text is accurate,
may reject it entirely, and assumes responsibility for the final report. These features support riskproportionate treatment of the function, but the principal relief requested here is a narrowly defined policy of
premarket enforcement discretion—not a categorical exemption based solely on the practice of medicine.9
The more persuasive basis for a limited policy is the final function's low-risk architecture: source images
remain available; the intended user is the responsible specialist; the unverified output is technically
sequestered; the radiologist controls finalization; any text error is reversible before downstream reliance; and
the system is subject to validation, monitoring, and controlled change management.
FDA CDS Guidance, supra note 2, at 5-7, 21-22.
6 21 C.F.R. § 892.2090(a), (b)(1)(ii)-(iv).
FD&C Act section 520(o)(4)(A), 21 U.S.C. section 360j(o)(4)(A).
U.S. Food & Drug Admin., Policy for Device Software Functions and Mobile Medical Applications, supra note 4, at 2, 13-14, 24-25.
9 FD&C Act section 1006, 21 U.S.C. section 396.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 5
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
V. FDA Should Evaluate the Final Deployed Function, Not Latent Foundation-Model
Capability in Isolation
Foundation models may possess broad, emergent, or multipurpose capabilities. A single model may support
research, retrospective quality review, private drafting, clinician-facing decision support, patient-facing
communication, or autonomous diagnosis. Those uses differ materially in intended user, degree of autonomy,
clinical reliance, reversibility, and consequence of error.
FDA's discussion paper identifies the appropriate unit of analysis: the final user-facing device, as configured
and intended for real-world deployment, rather than the foundation model standing alone or an isolated
subcomponent.10 That principle should remain central. A model's capacity to generate diagnostic language
should not automatically impose the same regulatory burden on every downstream use of that model.
Objective intent remains important. An upstream provider that specifically markets and supports a model for
autonomous clinical diagnosis cannot avoid device responsibilities merely by calling it a foundation model,
component, application programming interface, or development tool. But device status and regulatory burden
should not turn solely on latent capability or on what a downstream party could theoretically configure.
A. Fine-tuning is not a reliable regulatory bright line
The term "fine-tuning" does not identify a stable threshold of clinical risk. A downstream implementation
may be materially changed through model-weight adjustment, prompting, retrieval-augmented generation,
structured clinical context, post-processing, output constraints, confidence thresholds, report templates, model
routing, ensemble methods, interface design, access controls, or workflow safeguards. Conversely, a heavily
fine-tuned model may still be deployed autonomously in a high-risk setting.
Regulatory treatment should therefore follow the final intended clinical function and deployment architecture,
while assigning appropriate responsibilities to the upstream model provider, downstream developer,
deploying institution, and interpreting physician.
VI. Independent Verifiability and Access to the Primary Source Material Are Material Risk
Distinctions
Some medical devices generate measurements or primary representations that a physician cannot
independently verify at the point of use. A clinician ordinarily cannot directly observe intra-arterial pressure
without relying on the measurement chain. A radiologist relies on an imaging acquisition system to create the
primary representation of anatomy. Such devices need not be 100 percent accurate — no medical device is
held to that standard — but their users may have limited ability to detect certain errors independently, making
premarket evidence and controls especially important.
A private VLM pre-draft is different in kind. The VLM does not create the source CT, MRI, radiograph,
ultrasound examination, or other primary diagnostic image. It creates a secondary textual representation of
source material that remains fully available to the specialist radiologist. The radiologist may inspect the
images, review priors and clinical history, identify unsupported statements, correct laterality or location,
delete erroneous language, and discard the generated text completely.
FDA's proposed risk framework recognizes related distinctions. It identifies continuously supervised and
autonomous functions as meaningfully different, notes that certain measurement or signal-processing outputs
may be difficult for users to assess independently, and asks whether traceability to primary source
information, downstream safeguards, reversibility, time pressure, and specialist expertise should modify risk.11
FDA GenAI Discussion Paper, supra note 1, at 11 (describing evaluation of the final user-facing device as configured and intended for realworld deployment rather than the foundation model standing alone).
FDA GenAI Discussion Paper, supra note 1, at 5-9 (discussing supervised and autonomous activity, independent assessment, reversibility,
downstream safeguards, time pressure, traceability to primary source information, and specialist expertise).
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 6
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Independent specialist verification against the primary source images is a material risk control. The AIgenerated text remains private and non-final, and no model-generated statement can independently enter the
clinical care pathway before the radiologist reviews the source images and issues the final report.
VII. Specialist Review and Supervised Operation
A. Specialist Review Is a Meaningful Risk Control When It Is Designed to Be Meaningful
The phrase "human in the loop" is too broad to carry regulatory weight by itself. A generalist who receives a
subspecialty AI diagnosis after the output has already entered the care pathway is not equivalent to the
credentialed radiologist who has the complete imaging examination and exclusive authority to issue the
report.
Specialist supervision should be considered a material control only when the system is designed so that
review is substantive rather than nominal. That requires source access, sufficient time, freedom to reject the
output, prevention of downstream dissemination, affirmative finalization, appropriate user training, and
monitoring of whether the workflow actually supports independent review.
The radiologist’s independent clinical judgment operates together with appropriate product safeguards to
protect patients. It is one component of a layered safety architecture that also includes technical access
controls, local validation, human-factors design, model documentation, post-deployment monitoring, version
control, rollback capability, and institutional governance.
B. Supervised Practice Is an Element of the Competency Model the Agency Has Described
The discussion paper's competency analogy appropriately recognizes that clinical competence develops
through structured assessment and supervised practice before independent practice. In diagnostic radiology, a
useful — though necessarily imperfect — workflow analogue is the attending radiologist's review of a
trainee's draft interpretation before that draft is released for clinical reliance. A resident or fellow may
generate preliminary text that the attending reviews against the source images, corrects, replaces, and
ultimately assumes responsibility for upon signature. The preliminary interpretation may contain
discrepancies and may influence the attending's review; the relevant safeguards are meaningful source-image
review, supervision, and final specialist authorship.
The legal categories are different. A trainee's professional work product is not commercially distributed
software and is not an article subject to medical-device regulation. The analogy is therefore not offered to
establish equivalent legal status. It is offered to illustrate that the competency model contains a supervisedpractice stage between examination and independent operation. Models also present scale, reproducibility,
and correlated-failure risks that individual trainees do not; those differences support local validation, humanfactors evaluation, continuous performance monitoring, and controlled model changes.
The structural point stands nonetheless. If the Agency builds premarket evaluation on the model by which
clinicians are qualified, that model contains a stage between examination and independent practice. TMI and
RADIN respectfully suggest that the framework recognize a corresponding graduated structure, with the
evidence expected calibrated to the stage rather than uniformly to independent practice:
• Stage one — benchmarking only. Non-clinical evaluation, with no patient-facing deployment. Corresponds to
examination.
• Stage two — supervised operation. Deployment restricted to the RSG-PD conditions set out in Section VIII,
with architecturally enforced specialist review, sequestered output, and prospectively defined performance
monitoring. Corresponds to supervised practice. Evidence appropriate at entry: benchmarking together with
representative retrospective evaluation or prospective shadow deployment.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 7
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
• Stage three — independent or care-team-facing operation. Outputs that reach clinicians, the medical record, or
patients without specialist verification. Corresponds to independent practice, with clinical confirmation
proportionate to autonomy and consequence of error.
This is not a request to relax the standard applicable to autonomous diagnosis. It is a request that the
framework not require a function to satisfy the standard for the third stage as the price of admission to the
second. Requiring evidence appropriate to independent operation before permitting any qualifying supervised
operation could unnecessarily restrict an important setting for generating evidence about human-AI
performance and future development. The enforcement-discretion policy requested in Section XVIII is one
mechanism by which the Agency could give effect to that second stage; a streamlined pathway with defined
special controls would be another.
VIII. Proposed Conditions for Radiologist-Supervised Generative Pre-Drafting
FDA should consider premarket enforcement discretion only when every condition below is satisfied. These
conditions are deliberately restrictive so that a private drafting aid cannot become an unreviewed clinical
reporting channel by design or operational drift.
A. Qualified specialist user
The intended user must be a licensed physician qualified and credentialed to independently interpret the
relevant imaging examination.
A function intended for use by a patient, non-radiologist clinician, administrative user, or trainee without
responsible specialist supervision should not qualify merely because a radiologist might later review the case.
B. Mandatory Source-Image Access and Independent Interpretation
Before report finalization, the qualifying workflow must require the assigned radiologist to access the
complete imaging examination through an integrated or linked diagnostic viewer, together with relevant
clinical information, comparison examinations, and other source material ordinarily available for
interpretation. The software must not obscure, truncate, substitute for, or condition access to the source
images.
The qualifying workflow must not permit image-suppressed report finalization. Although software cannot
determine the quality of the radiologist's cognitive review, it can prevent a workflow in which a report is
finalized without source-image access.
C. Technical Sequestration and Clinical Non-Persistence of Unverified Output
The pre-draft must remain within the interpreting radiologist's controlled workspace. Before radiologist
finalization, it must not enter the legal medical record, appear in a clinician or patient portal, transmit through
HL7, FHIR, API, email, or other messaging as a clinical report, or otherwise become available for treatment
or disposition decisions.
The generated pre-draft and subsequent physician actions may be retained in access-controlled audit and
quality records for validation, safety monitoring, and controlled model improvement, subject to applicable
privacy, security, and data-governance requirements. Such retained information must not be displayed,
transmitted, or treated as a clinical report. No timer, default selection, inactivity state, or batch process may
convert unverified text into a clinical communication.
D. No automatic finalization or clinical action
The function must not autonomously sign or finalize a report; communicate a critical result; order a test or
treatment; change patient disposition; direct emergency or non-emergency care; assign a diagnosis to the
medical record; alter medication; or initiate any other patient-care action.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 8
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
A workflow timer, inactivity, default selection, or batch process must not convert the absence of editing into
authorization.
E. Affirmative radiologist control
The radiologist must affirmatively issue the final report and must be able to edit any statement, delete
findings, add findings, replace the impression, reject the entire pre-draft, or start from a blank report.
The interface should make complete rejection easy and should not penalize the user for beginning without AIgenerated text.
F. Meaningful independent interpretation and human-factors evaluation
The developer and deploying institution should evaluate whether timing, visual prominence, confidence
presentation, explanation design, and workflow sequence encourage premature closure or inappropriate
reliance.
FDA need not mandate a universal image-first or text-first interface. The final configured workflow should be
tested to determine whether radiologists can identify and correct characteristic model errors under realistic
workload and time pressure.
G. Local validation before clinical activation
The final configured function should be validated on data representative of intended modalities, body parts,
protocols, scanner types, patient populations, disease prevalence, clinical settings, report conventions, imagequality conditions, and clinically relevant subgroups.
Validation may include retrospective testing, prospective shadow deployment, blinded clinician review,
paired aided-versus-unaided assessment, and clinician adjudication. A randomized prospective clinical trial
should not be presumed necessary for every private pre-drafting function; evidence should be proportionate to
risk and intended use.12
H. Clinically meaningful performance measures
Evaluation should not rely only on generic text-similarity metrics. It should measure clinically significant
omissions and false additions, wrong laterality or location, unsupported conclusions, internal contradictions,
critical-finding performance, complete-rejection rates, material-correction rates, AI-induced errors, AIcorrected human errors, subgroup performance, and net final-report quality.
Where sensitivity and specificity are applicable, they should be assessed for the human-AI team and
compared with a realistic unaided workflow.
I. Continuous performance monitoring
The deploying organization should monitor clinically significant discrepancies, correction and rejection rates,
drift by modality, facility, scanner, population, and subgroup, unexpected output patterns, latency or system
failures, user complaints, and potential patient-safety events.
Prospectively defined thresholds should trigger investigation, restriction, rollback, or suspension.
J. Controlled model change management
Radiologist corrections may be collected and used, subject to privacy and data-governance requirements, for
offline error analysis, development, and retraining.
FDA GenAI Discussion Paper, supra note 1, at 14-18 (discussing non-clinical benchmarking, retrospective evaluation, shadow deployment,
clinician adjudication, prospective studies, comparators, and human-AI team performance).
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 9
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
A production model should not automatically alter its clinical behavior after every correction without
controlled evaluation. Safer lifecycle management is: supervised use → structured feedback → offline
development → validation of a defined version → controlled deployment → continued monitoring. FDA's
PCCP guidance provides relevant principles for planned changes, validation methodology, and impact
assessment.13
K. Auditability and cybersecurity
The system should retain appropriate records of the model and configuration version, generated pre-draft,
material physician edits, final report, timing, relevant source inputs, system failures, and corrective actions.
Access control, data integrity, cybersecurity, vendor-change notification, and incident-response procedures
should be proportionate to the deployment.
L. Clear exclusions
Enforcement discretion under the RSG-PD category should not apply to the extent an unverified AI-generated
diagnostic output reaches a treating clinician or patient, is represented as a preliminary or final clinical
interpretation, initiates treatment or patient disposition, autonomously communicates a critical result, or
otherwise becomes an independent source of diagnostic reliance before specialist verification.
Image-based worklist prioritization should be analyzed as a separate software function. The presence of a
separately regulated or appropriately authorized prioritization function within the same platform should not
disqualify an otherwise qualifying RSG-PD function from enforcement discretion, provided that the pre-draft
itself remains sequestered and all applicable requirements governing the prioritization function are
independently satisfied.
IX. Automation Bias Must Be Evaluated Against the Real-World Baseline of Human
Interpretation
AI-generated text can create anchoring, automation bias, confirmation bias, or premature diagnostic closure.
Incorrect AI advice has been shown to influence radiologists, and human-factors interventions can change the
degree of influence.14 Explanation design can also affect diagnostic performance and trust, including in ways
users may not consciously recognize.15 The effect of AI assistance is heterogeneous across radiologists and
tasks, and AI error is a major predictor of whether assistance helps or harms performance.16
Those risks should be measured and mitigated against the actual baseline of human interpretation rather than
an assumption of error-free unaided performance.
Unaided radiology is also subject to perceptual and cognitive influences, including expectation effects,
anchoring, satisfaction of search, prevalence effects, fatigue, interruptions, time pressure, incomplete clinical
information, and sequential dependence on recently reviewed cases. Experimental research has shown that
radiologists' judgments of simulated lesions on a current radiograph can be pulled toward stimuli seen in the
preceding one or two radiographs. This supports the narrower principle that recent visual experience can
influence current interpretation.17
U.S. Food & Drug Admin., Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial IntelligenceEnabled Device Software Functions 4-12 (Aug. 2025), https://www.fda.gov/media/166704/download [hereinafter FDA PCCP Guidance].
Michael H. Bernstein et al., Can Incorrect Artificial Intelligence (AI) Results Impact Radiologists, and If So, What Can We Do About It? A
Multi-Reader Pilot Study of Lung Cancer Detection with Chest Radiography, 33 Eur. Radiol. 8263, 8263-69 (2023), doi:10.1007/s00330-02309747-1.
Drew Prinster et al., Care to Explain? AI Explanation Types Differentially Impact Chest Radiograph Diagnostic Performance and Physician
Trust in AI, Radiology 313(2), e233261 (2024), doi:10.1148/radiol.233261, PMID 39560483.
Feiyang Yu et al., Heterogeneity and Predictors of the Effects of AI Assistance on Radiologists, 30 Nature Medicine 837, 837-49 (2024),
doi:10.1038/s41591-024-02850-w.
Mauro Manassi et al., Serial Dependence in the Perceptual Judgments of Radiologists, 6 Cognitive Research: Principles and Implications 65
(2021), doi:10.1186/s41235-021-00331-z, PMID 34648124.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 10
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
AI-related anchoring and overreliance should therefore be evaluated as additional human-factors risks
superimposed on an already imperfect diagnostic process, rather than against a hypothetical baseline of errorfree unaided interpretation. The appropriate comparison is the final AI-assisted workflow against realistic
unaided practice.
The assessment should measure both directions of effect: errors introduced or reinforced by the AI; errors
detected or prevented with AI assistance; changes in sensitivity, specificity, completeness, consistency, and
interpretation time; effects under fatigue and workload; and variability among users. Research evaluating
VLM-generated radiology reports has found clinically significant errors in both AI-generated and humanauthored reports, further supporting evaluation of the collaborative system rather than treating either
participant as inherently error-free.18
Both directions of effect matter to a benefit-risk assessment, and the second deserves explicit attention
because it is easily overlooked in a framework organized around the containment of error. Most clinically
significant radiology reporting errors are omissions rather than misinterpretations, and the perceptual and
cognitive influences described above — satisfaction of search, fatigue, interruption, and time pressure —
operate principally by causing findings to go unreported. A structured pre-draft may reduce those omissions
through mechanisms that do not depend on the model “diagnosing” anything: it presents an enumerated set of
observations that the radiologist must confirm or reject against the source images, so that a reader who has
already identified one abnormality is prompted to address the others; it surfaces comparison examinations and
interval change that might otherwise not be reviewed; and it does so most consistently at exactly the points in
a shift where unaided human performance is weakest. In each case the radiologist remains the one who sees,
interprets, and decides. The pre-draft functions as a structured second look that the same specialist
adjudicates, in the manner that double reading in screening mammography exists to catch what a single
reading misses.
TMI and RADIN do not assert that this benefit is established; they assert that it is measurable, and that the
human-AI team evaluation described in Section X and Questions 14 and 15 is the instrument for measuring it.
A framework that quantifies only the errors assistance introduces, and never the errors it prevents, cannot
support the benefit-risk determination that the FD&C Act requires. The evaluation of a qualifying RSG-PD
function should therefore be designed to capture omissions avoided, findings addressed that would otherwise
have gone unreported, and completeness of the final signed report, alongside the automation-bias measures
described above.
Potential human-factors safeguards include:
• clear identification of unverified AI-generated text;
• a readily available blank-report option;
• no automatic or passive acceptance;
• training regarding known failure modes and appropriate skepticism;
• interface testing of draft timing, confidence presentation, and explanation design;
• monitoring of AI-induced errors separately from pre-existing human errors;
• periodic adjudication of representative final reports;
• assessment of individual and subgroup differences in reliance;
• periodic evaluation of unaided performance where appropriate; and
• modification or suspension of workflows that produce unacceptable anchoring or error propagation.
X. Human-AI Team Performance Should Be the Principal Performance Unit
For an RSG-PD function, the intended clinical product is not the VLM acting independently. It is the
combined workflow in which the VLM produces provisional text and the radiologist produces the clinical
Ryutaro Tanno et al., Collaboration Between Clinicians and Vision-Language Models in Radiology Report Generation, 31 Nature Medicine
599, 599-608 (2025), doi:10.1038/s41591-024-03302-1.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 11
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
interpretation. Requiring the VLM alone to meet the standard expected of an autonomous diagnostician
would test a function the product is not intended to perform.
FDA's discussion paper expressly asks when human-AI team performance should be evaluated instead of
device-alone performance and whether the appropriate comparator may be the care likely to occur without the
device, including unaided clinical judgment.19 International Good Machine Learning Practice principles
likewise emphasize performance of the human-AI team when a human is part of the intended use and
monitoring of deployed models and retraining risk.20
The principal questions should be:
• Does use of the pre-draft preserve or improve the accuracy and completeness of the final report?
• Does it create clinically significant new errors, and can users reliably detect characteristic failures?
• Does it reduce or increase variability among radiologists?
• Does it improve timeliness without compromising independent interpretation?
• Does performance remain acceptable across sites, equipment, populations, subgroups, and clinically relevant
abnormalities?
• Does the interface create inappropriate reliance, and do proposed mitigations work?
• Can adverse trends be detected, investigated, and reversed promptly?
XI. Downstream Clinical Reliance Should Be the Principal Regulatory Boundary
The most practical boundary is whether an AI-generated interpretation can be relied upon before specialist
verification. When the output is private, reversible, sequestered, and reviewed against complete source
images by the responsible specialist, multiple safeguards intervene before patient care can be affected. When
the output reaches a clinician or patient first, the AI interpretation has entered the clinical decision chain.
Protected specialist-supervised pre-drafting Higher-risk clinical-output pathway
Images → private AI pre-draft → interpreting
Images → AI diagnostic or preliminary output
radiologist independently reviews the source
→ clinician or patient receives or may rely on
images → radiologist edits, rejects, replaces, or
the output → care decision may occur before
approves the text → radiologist signs the final
specialist verification → radiologist review may
report → clinician receives only the radiologistoccur later or not at all.
issued report.
Proposed treatment: narrowly bounded premarket Proposed treatment: appropriate FDA premarket
enforcement discretion, with validation, monitoring, authorization and controls proportionate to autonomy,
auditability, and change controls. clinical reliance, and consequence of error.
FDA should apply appropriate premarket authorization when an AI function substitutes for specialist
interpretation; reaches a user who cannot independently evaluate the source images; is available for
reliance before specialist confirmation; autonomously initiates or directs care; communicates a critical
result; changes treatment priority; or is represented as a clinically actionable interpretation.
XII. Local Clinical Validation Is a Safety Strength, Not a Regulatory Loophole
Imaging AI performance can vary with scanner manufacturer, acquisition protocol, image quality, patient
demographics, disease prevalence, referral pattern, comparison-study availability, institutional terminology,
FDA GenAI Discussion Paper, supra note 1, at 16, 18 (Questions 14 and 15).
U.S. Food & Drug Admin., Health Canada & Medicines and Healthcare products Regulatory Agency, Good Machine Learning Practice for
Medical Device Development: Guiding Principles, principles 7 and 10 (Oct. 2021), https://www.fda.gov/media/153486/download.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 12
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
and clinical setting. A model that performs adequately in one environment may not perform adequately in
another. Local validation is therefore a necessary safety layer.
A proportionate pathway should allow qualified institutions to validate the final configured system on
representative data, begin in retrospective or shadow mode, establish prospective thresholds, deploy gradually
by modality or use case, monitor discrepancies, maintain rollback capability, and share material safety
findings with relevant developers and regulators.
An Institutional Review Board ("IRB") may be required when a particular activity constitutes human-subject
research or a regulated clinical investigation. An IRB is not, however, a general substitute for medical-device
regulation or routine clinical AI governance. Applicable requirements under 21 C.F.R. parts 50, 56, and 812
should be assessed for the specific activity, while routine quality-improvement deployment should be
governed through appropriate institutional validation, clinical oversight, privacy, information security, and
safety processes.21
XIII. Radiologist Feedback Can Improve Safety Through Controlled Learning
Specialist-supervised pre-drafting can generate unusually valuable real-world evidence. Each accepted
statement, correction, deletion, addition, or complete rejection can identify an omitted finding, unsupported
finding, wrong laterality, wrong location, inappropriate severity, failure to use a comparison study, failure to
answer the clinical question, or unsuitable recommendation.
Subject to appropriate authorization, privacy, data governance, and quality controls, this feedback can support
error-taxonomy development, hard-negative case identification, subgroup analysis, offline retraining, test-set
expansion, improved prompts and guardrails, and more clinically meaningful future validation.
The production system should not silently or autonomously learn from each case. The appropriate cycle is
controlled: supervised use → structured feedback → offline analysis and development → validation of a
defined new version → controlled release → post-deployment monitoring. This preserves the scientific
benefit of learning from practice without creating an unbounded, continuously changing clinical model.
XIV. Stronger Post-Deployment Monitoring Can Support a Proportionate Premarket
Burden
Open-ended generative systems may be difficult to characterize exhaustively before deployment. FDA's
discussion paper asks whether greater postmarket monitoring can justify acceptance of greater premarket
uncertainty in appropriate circumstances and discusses periodic benchmarking, sample-based clinician
review, drift monitoring, and shared ecosystem responsibilities.22
A private specialist-supervised pre-drafting function is a strong candidate for this approach because the
unverified output is intercepted before downstream reliance, the specialist creates immediate feedback,
deployment can be limited incrementally, performance can be monitored in real time, and the function can be
suspended or rolled back promptly.
Postmarket monitoring is not a substitute for adequate pre-deployment evidence. It is an additional control
that can make the total evidence package proportionate to the actual risk of the intended function.
Any reduced premarket burden should depend on technically enforced RSG-PD conditions, a prespecified
monitoring plan with performance thresholds and defined action triggers, correction-level auditability,
rollback capability, and mandatory restriction or suspension when a threshold is breached.
FDA GenAI Discussion Paper, supra note 1, at 15-16 & n.20 (noting that requirements under 21 C.F.R. parts 50, 56, and 812 may apply to
clinical-confirmation activities depending on the circumstances).
FDA GenAI Discussion Paper, supra note 1, at 19-21 (discussing postmarket monitoring, periodic benchmarking, sample-based clinician
review, drift, shared responsibility, and questions concerning premarket uncertainty).
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 13
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
XV. Responsibilities Should Be Shared Without Diffusing Accountability
A. Upstream foundation-model or technology provider
• Identify model versions and material changes;
• provide safety-relevant documentation, known limitations, and intended or prohibited uses;
• support version pinning and change notification where feasible;
• maintain appropriate cybersecurity and data-integrity controls; and
• avoid unsupported clinical marketing claims.
B. Downstream developer or integrator
• Define the intended use and final deployment architecture;
• implement technical sequestration and workflow controls;
• validate the final configured function;
• evaluate human factors;
• establish monitoring and rollback thresholds;
• manage model, prompt, retrieval, and post-processing changes; and
• maintain auditability and incident response.
C. Deploying healthcare organization
• Authorize qualified users and train them on limitations;
• conduct local validation and phased deployment;
• monitor real-world performance and investigate safety events;
• restrict or suspend problematic functions;
• maintain governance, privacy, and security controls; and
• ensure unverified output cannot bypass the intended clinical workflow.
D. Interpreting radiologist
• Independently review the source images and relevant clinical information;
• treat the pre-draft as unverified work product;
• correct, reject, or replace inaccurate text;
• issue the final report based on professional judgment; and
• report material or recurrent system failures through the established quality process.
FDA's proposed voluntary Foundation Model Device Master File concept may improve confidential
information sharing and reduce duplication. Importantly, the discussion paper states that submission of such a
master file would not authorize the underlying foundation model for a device intended use and that the
downstream sponsor would remain responsible for its final device.23 That approach is more appropriate than
treating every model with potential diagnostic capability as a finished diagnostic device for every downstream
use.
XVI. Responses to FDA Discussion Questions
TMI and RADIN address all 26 questions presented in the FDA discussion paper to provide a complete
record regarding radiologist-supervised image-to-report pre-drafting. These responses should be read together
with the analysis and authorities presented in Sections IV through XV.
Several questions concern functions that fall outside the proposed Radiologist-Supervised Generative PreDrafting, or RSG-PD, category, including multi-turn patient conversations, autonomous care escalation, and
FDA GenAI Discussion Paper, supra note 1, at 22 (describing a potential voluntary Foundation Model Device Master File and noting that
such a submission would not itself authorize the underlying model for a device intended use).
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 14
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
agentic AI. For those questions, TMI and RADIN identify the distinction and provide general
recommendations.
The responses to Questions 7 through 17 are provided without waiving TMI and RADIN's position that a
qualifying RSG-PD function should be subject to premarket enforcement discretion. If FDA determines that a
particular implementation requires premarket review, the nature and amount of evidence should remain
proportionate to the function's actual intended use, deployment safeguards, and residual risk.
A. Considerations for the Assessment of Risk
Question 1: Does the two-axis framework capture the relevant dimensions of risk?
The proposed axes of independent device activity and the consequences of relying on an incorrect output are
useful but incomplete.
FDA should also consider:
• traceability to the primary source information;
• the user's expertise and ability to independently evaluate that source information;
• technical barriers preventing downstream dissemination;
• reversibility before patient-care impact;
• time pressure within the deployment setting;
• whether the output can initiate or direct a clinical action;
• whether specialist review is mandatory;
• the adequacy of post-deployment monitoring and rollback controls; and
• whether the model or surrounding configuration can change without controlled validation.
These factors materially distinguish a private RSG-PD function from autonomous or externally relied-upon
diagnostic software. In a qualifying RSG-PD workflow, the radiologist has the complete source images, the
AI-generated text remains private and reversible, and the output cannot independently reach or influence
patient care.
A concrete illustration may assist the Agency. Consider two deployments of the identical model, producing
identical text from the identical chest radiograph. In the first, the draft appears in the interpreting radiologist’s
workspace beside the images, remains private and non-final, and cannot be released as a clinical report
without the radiologist’s independent source-image review and affirmative finalization. In the second, the
same text is written to the electronic health record as a preliminary result and transmitted to the overnight
emergency physician before specialist verification. These deployments differ materially in clinical reliance
and the consequences of error, even though the underlying model, images, and generated text are identical.
Explicitly identifying output disposition, reversibility, source-image access, and mandatory review would
help ensure that the framework consistently captures those differences.
One of those dimensions warrants particular emphasis. The activity axis distinguishes continuous healthcare
professional supervision from autonomous operation, but supervision is not binary. Supervision that a
workflow permits the user to bypass is not equivalent to supervision supported by mandatory workflow
controls. TMI and RADIN respectfully urge the Agency to distinguish architecturally enforced review
requirements from advisory review. Required source-image access and affirmative specialist sign-off are
design controls that can be inspected in system configuration and audit logs; an instruction to review, without
corresponding workflow controls, remains an expectation about user behavior.
Question 2: How should FDA distinguish non-directive information from action-directing
information?
Directiveness should be evaluated functionally and in the context of actual deployment, rather than
determined solely by the wording of the output.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 15
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Relevant considerations include:
• who receives the output;
• the output's specificity, personalization, urgency, and expressed certainty;
• whether it assigns a diagnosis or proposes treatment, triage, disposition, or follow-up;
• whether it enters an EHR, alerting system, ordering system, patient portal, or clinical communication channel;
• whether it can automatically initiate an order, alert, communication, or workflow change;
• whether the recipient can independently evaluate the underlying source information; and
• whether specialist verification is required before the output can affect care.
A disclaimer should not convert an otherwise action-directing output into a non-directive one. Conversely, the
presence of diagnostic language should not automatically make an output action-directing when that output is
private, non-final, inaccessible to downstream users, and incapable of affecting care before specialist
verification.
A qualifying RSG-PD function therefore falls toward the lower end of the directiveness continuum. Its output
becomes clinically operative only when the radiologist independently reviews the images and issues the final
report.
Question 3: When does delivering clinical information directly to patients create different or higher
risk?
Unverified GenAI-generated clinical information may present greater risk when delivered directly to a patient
who may not have the expertise necessary to assess its accuracy, limitations, or clinical significance. The risk
is particularly important when the output includes a diagnosis, prognosis, treatment recommendation, or
instruction about whether to seek or defer care.
Patient-facing status should not automatically make every function high risk. Administrative information,
general educational material, and access to a clinician-finalized report are materially different from an
unverified AI-generated diagnostic interpretation.
For a qualifying RSG-PD function, the AI pre-draft should not be available to the patient or to any clinician
other than the interpreting radiologist. Technical controls should prevent it from appearing in the patient
portal, EHR, clinician inbox, API output, or any other downstream communication channel.
The patient and treating clinician should receive only the report issued by the radiologist after specialist
review. This safeguard applies to unverified AI-generated work product and should not restrict a patient's
appropriate access to the finalized radiology report.
Question 4: How should the distinction between generalist and specialist physicians affect risk
assessment?
Specialist expertise should be a meaningful risk modifier when the specialist:
• has access to the original source information;
• is practicing within the relevant clinical specialty;
• possesses the training necessary to evaluate the output independently;
• exclusively controls finalization;
• can readily reject or replace the AI output; and
• receives appropriate training regarding the function's limitations.
A radiologist independently interpreting a complete imaging examination is differently situated from a
generalist clinician receiving an unverified subspecialty diagnosis. The radiologist has both the source images
and the specialty expertise necessary to identify unsupported or incorrect statements.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 16
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Specialist status alone should not eliminate oversight. Its significance arises when it operates together with
source-image access, output sequestration, affirmative physician finalization, technical safeguards, local
validation, and performance monitoring.
Question 5: How should risk be evaluated across multi-turn conversational trajectories?
Risk for a conversational GenAI-enabled function should be assessed across complete and realistically
foreseeable conversational trajectories, rather than one isolated response at a time.
Testing should determine whether a function that begins within a non-directive scope can migrate toward
diagnosis, treatment recommendations, triage instructions, or other action-directing behavior. The risk
assessment should address reasonably foreseeable conversational behavior, including migration beyond the
intended scope. The intended-use statement should clearly identify the supported clinical tasks, users, and
limits; an out-of-scope capability should not, by itself, redefine intended use.
Appropriate safeguards may include:
• hard scope boundaries;
• restrictions on the types of outputs that can be generated;
• refusal behavior when the conversation moves outside the intended use;
• testing of adversarial and emotionally framed prompts;
• monitoring of cumulative conversational context; and
• clear escalation to a qualified human when appropriate.
Question 6: How should under-escalation and over-escalation be evaluated?
Under-escalation and over-escalation should be measured separately because they produce different types of
harm and may not be appropriately represented by a single aggregate accuracy measure.
Evaluation should consider:
• the clinical severity of missed escalation;
• the consequences of unnecessary escalation;
• disease prevalence;
• urgency;
• available alternative safeguards;
• the intended user;
• the availability of immediate professional review; and
• the effect of each error direction on patients and healthcare resources.
Acceptance criteria should be specific to the clinical context. A missed emergency escalation may warrant a
substantially lower tolerance than an unnecessary recommendation for non-emergency follow-up.
B. Competency-Based Premarket Evaluation
Question 7: Is benchmarking followed by clinical confirmation an appropriate evaluation framework?
For GenAI-enabled functions subject to premarket review, device benchmarking followed by riskproportionate clinical confirmation is a useful general framework.
The evaluation should focus on the final user-facing function as configured and intended to be deployed,
rather than on the foundation model acting in isolation.
For an RSG-PD function, the intended system is the combined workflow:
VLM-generated private pre-draft → radiologist review of the source images → editing,
rejection, or replacement of the pre-draft → radiologist-issued final report.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 17
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Clinical confirmation should evaluate the performance of that workflow. It should not require the VLM
operating alone to satisfy the standard applicable to an autonomous diagnostician when autonomous use is
neither intended nor technically permitted.
Appropriate methods may include representative retrospective testing, prospective shadow deployment,
specialist adjudication, paired aided-versus-unaided evaluation, and human-factors testing. The rigor of
confirmation should reflect residual risk after accounting for specialist verification, output sequestration,
reversibility, and downstream safeguards.
The commenters respectfully refer the Agency to Section VII.B above, which proposes that the competency
framework recognize a supervised-operation stage between non-clinical benchmarking and independent
operation, with the evidence expected at each stage calibrated accordingly.
Question 8: How should the risk framework determine the level of premarket evidence?
The amount and rigor of premarket evidence should be based on the residual risk of the final deployed
function after application of its actual safeguards. It should not be based solely on the latent capabilities of the
underlying foundation model or on a hypothetical autonomous use that is outside the intended use.
Relevant considerations include:
• who receives the output;
• whether the user can independently evaluate the source information;
• whether specialist review is mandatory;
• whether the output remains technically sequestered until approval by a specialist;
• whether it is reversible before clinical reliance;
• whether it can initiate a clinical action;
• whether the system can bypass the specialist; and
• whether monitoring and rollback controls are in place.
A private RSG-PD function has limited independent clinical activity because it does not issue a report,
communicate with a clinician or patient, trigger treatment or triage, or bypass radiologist review.
Greater evidence should be required as a function moves toward direct delivery to clinicians or patients,
action-directing behavior, time-critical use, reduced opportunity for specialist verification, or autonomous
clinical action.
Question 9: Is the proposed benchmarking structure adequate, and what additional elements are
needed?
FDA's proposed benchmarking structure is a useful starting point, but multimodal radiology VLMs require
additional evaluation tailored to the relationship between the source images and the generated text.
Testing should assess:
• whether clinically material statements are supported by the source images;
• clinically significant omissions and unsupported additions;
• laterality, location, severity, number, and measurement errors;
• complete-study ingestion and coverage;
• comparison with prior examinations;
• consistency between the findings and impression;
• clinically meaningful variability across repeated outputs;
• generalizability across modalities, protocols, scanners, institutions, and populations;
• subgroup performance;
• workflow containment; and
• human-AI team performance.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 18
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Generic text-similarity metrics should not serve as the principal measure of clinical performance. Two reports
may differ significantly in wording while remaining clinically equivalent, or appear linguistically similar
while containing a material diagnostic error.
For supervised pre-drafting, testing should also measure material correction rates, complete-draft rejection,
AI-induced errors, radiologist-corrected AI errors, verification time, report completeness, and final-report
quality compared with realistic unaided interpretation.
The final configured system should also be tested to confirm that an unverified pre-draft cannot be
automatically finalized, transmitted, or used to initiate clinical action.
Two additional benchmarking elements warrant separate treatment because they are not captured by accuracy
measures. The first is fabricated specificity: whether the output asserts measurements, comparisons, technique
parameters, or clinical history that are not derivable from the inputs supplied. This is the report-generation
form of confabulation, and it is distinct from diagnostic error because the fabricated assertion may be
diagnostically plausible and nonetheless wholly invented.
The second is direct measurement of whether specialist review is functioning as designed. Where a function
is deployed under the conditions proposed in Section VIII, the sponsor and deploying organization can
measure verbatim-adoption rates, edit-distance distributions, per-reader variation in adoption, and blinded
independent re-read of a random sample of signed reports stratified by whether the pre-draft was adopted
verbatim or edited. Each can be evaluated without communicating unverified output to downstream clinicians
or patients. TMI and RADIN specifically caution against approaches that would seed known errors into live
clinical pre-drafts in order to test reviewer vigilance; retrospective blinded re-read can assess reviewer
performance without deliberately introducing error into patient care.
These measures are surveillance signals, not substitutes for independent clinical adjudication, and they do not
by themselves establish that a radiologist performed an adequate cognitive review. They should be interpreted
together with blinded review of final reports and independent measures of clinical accuracy.
Question 10: How should benchmark contamination, saturation, and representativeness be
addressed?
Benchmark performance should be treated as predictive of real-world performance only when the benchmark
is clinically relevant, appropriately independent from model development, and representative of the intended
deployment environment.
Reasonable controls may include:
• sequestered datasets;
• temporal or institutional holdouts;
• documented data provenance;
• screening for duplicate and near-duplicate cases;
• prespecified endpoints and scoring criteria;
• a locked model and configuration before testing;
• independent specialist adjudication;
• repeated-run testing;
• subgroup analysis; and
• confirmation using representative real-world data or prospective shadow deployment.
Public benchmarks should not automatically be presumed superior. Public availability may increase the risk
of training-data contamination, repeated optimization, and benchmark saturation.
Sponsor-developed benchmarks should remain permissible where no adequate public benchmark exists,
provided the sponsor documents the case-selection method, data provenance, scoring methodology,
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 19
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
acceptance criteria, adjudication process, and known limitations. Sponsor-developed benchmarks should not
ordinarily constitute the sole evidence supporting a material clinical claim.
Where a downstream developer lacks complete visibility into an upstream foundation model's training data,
FDA should permit reasonable demonstrations of benchmark independence rather than require impossible
certainty that no benchmark case was ever encountered during upstream training.
Question 11: How should the method of clinical confirmation be selected?
For a private radiologist pre-draft, appropriate methods may include:
• retrospective local validation;
• prospective shadow deployment;
• blinded specialist review;
• paired aided-versus-unaided comparison;
• adjudication of clinically significant discrepancies;
• human-factors evaluation;
• phased deployment; and
• prospective performance monitoring.
The confirmation method should be proportional to the function's autonomy, clinical reliance, novelty,
intended users, deployment environment, and consequences of error.
A prospective randomized clinical study should not automatically be required for every private, specialistsupervised pre-drafting function. Lower-risk functions may be adequately evaluated through representative
retrospective data, shadow deployment, and specialist adjudication.
More rigorous prospective evidence may be appropriate when the AI output reaches treating clinicians or
patients, operates in a time-critical environment, initiates action, replaces specialist interpretation, or
otherwise presents a direct pathway from model output to patient care.
Question 12: How should sponsors obtain statistically meaningful performance measurements?
Sponsors should prespecify clinically meaningful endpoints, the unit of analysis, sample-size assumptions,
acceptance criteria, and the statistical methods used to account for variation among cases, readers,
institutions, and repeated model outputs.
Evaluation should report:
• point estimates and confidence intervals;
• clinically significant error rates;
• reader- and site-level variability;
• performance across modalities and use cases;
• subgroup estimates;
• repeated-run variability; and
• performance under representative real-world conditions.
When the same radiologists or cases contribute multiple observations, the analysis should account for
clustering and repeated measures.
Synthetic inputs and real-patient data should not automatically be pooled into a single performance estimate.
Pooling is appropriate only when the sponsor establishes that the data sources measure the same clinical
construct and adequately addresses differences in distribution. Otherwise, synthetic and real-world results
should be reported separately, with real-patient data serving as the principal evidence of clinical performance.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 20
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Question 13: When is synthetic data appropriate or inadequate?
Synthetic data may be useful for:
• stress testing;
• rare or dangerous edge cases;
• controlled alteration of specific imaging characteristics;
• testing system boundaries;
• supplementing underrepresented scenarios;
• evaluating robustness; and
• reducing unnecessary exposure of identifiable patient information.
Synthetic data should not replace representative real-patient imaging data when validating the clinical
performance of an image-to-report function.
It may be particularly inadequate when subtle pathology, acquisition artifacts, disease prevalence,
demographic characteristics, or institutional variation cannot be reproduced faithfully. There is also a risk that
synthetic data generated by a related model class may reproduce the same biases or blind spots that the
evaluation is intended to detect.
Safeguards should include independent clinical review, documentation of the generation method, comparison
with real-world distributions, subgroup-specific analysis, use of multiple data sources where feasible, and
confirmation of material findings on representative real-patient data.
Question 14: How should comparators and acceptance criteria be selected for open-ended outputs?
For open-ended radiology outputs, a qualified specialist panel or adjudicated clinical reference standard may
be appropriate when a single definitive reference report does not exist.
Acceptance criteria should focus on clinically significant agreement rather than exact wording. Relevant
measures include:
• missed or unsupported findings;
• laterality and location errors;
• incorrect measurements;
• inappropriate diagnostic conclusions;
• findings-impression inconsistency;
• clinically significant recommendations; and
• effects on the final radiologist-issued report.
For supervised pre-drafting, the human-AI team should be the principal evaluation unit. The relevant question
is whether the radiologist using the final configured function produces reports that are at least as safe and
clinically accurate as the realistic unaided workflow.
Device-alone performance remains useful for characterizing failure modes and defining safeguards. It should
not be the sole clinical standard when autonomous interpretation is outside the intended use and technically
prohibited.
Question 15: Should performance be compared with the care likely to occur without the device?
Yes. The comparator should reflect the actual care, technology, and workflow likely to exist in the function's
absence.
For an RSG-PD function, the principal comparator should ordinarily be realistic unaided radiologist
reporting. The evaluation should account for normal case complexity, workload, fatigue, interruptions, time
pressure, and ordinary human error.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 21
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
The appropriate comparator may differ for other functions. For example, delayed specialist review, generalist
interpretation, or no intervention may be appropriate where those conditions represent the actual alternative
available to the intended population.
For access and timeliness claims, the relevant comparator may differ from the comparator used to assess
diagnostic accuracy. In many overnight, rural, critical-access, and off-hours imaging settings, delayed
specialist review may be the actual alternative rather than a hypothetical one. The realistic alternative to an
assisted interpretation may be an interpretation rendered hours later or a treating clinician's provisional
assessment in the interim.
TMI and RADIN respectfully suggest that a sponsor be permitted to justify a delayed-specialist-review or
unaided-judgment comparator by documenting the actual service pattern at the intended deployment sites —
turnaround-time distributions, hours of specialist availability, and the interpreting arrangements in place
absent the function — rather than by assertion. That evidence is routinely captured in radiology operational
systems and can be produced without a prospective study.
The comparison should measure both directions of effect:
• errors introduced or reinforced by AI assistance; and
• errors prevented, identified, or corrected with AI assistance.
A GenAI-enabled function should be evaluated against the real clinical baseline, not against a hypothetical
error-free physician or an autonomous use the product is not intended to perform.
Question 16: What role should independent third parties play?
Independent third parties may contribute meaningfully to selected aspects of competency-based evaluation,
but participation should be risk-proportionate rather than universally mandatory.
Appropriate roles may include:
• maintaining sequestered datasets;
• conducting independent specialist adjudication;
• auditing benchmark methodology;
• validating scoring rubrics; and
• confirming performance for higher-risk or autonomous functions.
For a lower-risk RSG-PD function, adequate independence may also be achieved through radiologists who
were not involved in model development, blinded adjudication, external review of a representative sample, or
institution-specific validation under formal AI governance.
A universal third-party mandate could create unnecessary cost, delay, limited testing capacity, and barriers to
entry. FDA should preserve multiple valid means of achieving independent evaluation and should avoid
allowing a small number of organizations to control essential datasets or certification pathways.
Third-party involvement should supplement, not replace, the responsibilities of the developer, downstream
integrator, and deploying healthcare organization.
Question 17: How should the competency-based framework be adapted for multimodal VLMs?
A competency-based framework can be applied to multimodal VLMs, but it must address the relationship
among image inputs, clinical context, and generated language.
Evaluation should include:
• cross-modal grounding of generated statements;
• complete-study ingestion and coverage;
• clinically significant omissions and hallucinations;
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 22
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
• laterality, localization, measurement, and interval change;
• appropriate use of clinical history and prior examinations;
• findings-impression consistency;
• output variability and reproducibility;
• generalizability across populations and imaging environments;
• human-factors effects; and
• sensitivity to changes in the model, prompts, retrieval sources, templates, guardrails, post-processing, and
interface.
The final configured function — not merely the underlying foundation model — should be the primary unit
of evaluation.
For an RSG-PD function, the central performance question should be whether the radiologist-plus-VLM
workflow produces final reports that are at least as safe and clinically accurate as realistic unaided reporting,
without permitting unverified AI output to reach or influence downstream care independently.
Standalone model performance may still be evaluated to identify failure modes, but it should not replace
human-AI team performance as the principal evaluation of the intended supervised function.
C. Postmarket Monitoring for GenAI-Enabled Functions
Question 18: When may stronger postmarket monitoring justify greater premarket uncertainty?
Greater reliance on postmarket monitoring may be appropriate when the function:
• has limited independent clinical activity;
• operates under continuous specialist supervision;
• produces an output that remains reversible before clinical reliance;
• cannot directly reach a patient or treating clinician;
• is technically contained from initiating clinical action;
• can be monitored using meaningful real-world performance measures;
• can be promptly disabled or rolled back; and
• is deployed incrementally under defined thresholds.
A qualifying RSG-PD function possesses many of these characteristics because the radiologist intercepts the
output before it enters the patient-care pathway and creates immediate feedback through edits, rejection, or
replacement of the text.
Greater premarket uncertainty would be less appropriate for autonomous diagnosis, direct-to-patient or directto-clinician output, irreversible or high-consequence action, time-critical use without meaningful review, or
functions for which performance cannot be reliably monitored after deployment.
Question 19: What postmarket monitoring methods, cadence, and triggering events are appropriate?
FDA's proposed approaches — periodic re-benchmarking, sample-based clinician review, and performancedegradation monitoring — are appropriate components of a postmarket program.
Additional measures for RSG-PD functions may include:
• clinically significant correction rates;
• complete-draft rejection rates;
• unsupported additions and significant omissions;
• laterality and measurement errors;
• subgroup performance;
• AI-induced versus AI-corrected radiologist errors;
• complaints and safety events;
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 23
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
• workflow containment failures; and
• performance by modality, facility, scanner, or protocol.
Monitoring cadence should be proportionate to risk, utilization, and the maturity of the function.
Triggering events should include:
• a foundation-model or fine-tuned-model update;
• material prompt, retrieval, guardrail, or post-processing changes;
• workflow or interface changes;
• new modalities or intended populations;
• significant changes in scanners or acquisition protocols;
• a material change in case mix or disease prevalence;
• onboarding of a new facility or materially different reporting environment;
• substantial turnover or change in the supervising radiologist pool;
• a sustained unexplained change in verbatim-adoption, editing, or rejection patterns;
• a safety incident;
• evidence of drift; or
• breach of a prospectively defined performance threshold.
Question 20: Can machine-based supervisory agents facilitate postmarket monitoring?
Machine-based supervisory agents may help screen large volumes of data, identify unusual outputs, detect
drift, and flag cases for human review. They may be particularly useful for identifying laterality
inconsistencies, unsupported language, findings-impression contradictions, abnormal correction rates, or
deviations from established output patterns.
A supervisory agent should not serve as the sole postmarket safeguard.
Its own performance should be independently evaluated, including:
• sensitivity and specificity for the monitored failure modes;
• reliability across model versions and clinical settings;
• susceptibility to drift;
• false-positive burden;
• failure visibility;
• auditability; and
• escalation to qualified human reviewers.
FDA should also consider correlated-failure risk. A supervisory agent built on the same model family,
training data, or architecture as the monitored function may share its blind spots.
Machine-based monitoring should therefore supplement periodic human adjudication, not replace it.
Question 21: How should postmarket responsibilities be distributed without diffusing manufacturer
accountability?
The developer or manufacturer of the final configured function should remain accountable for its safety,
performance, change management, and postmarket monitoring.
Other stakeholders may provide important complementary functions:
• healthcare institutions can conduct local validation, monitor site-specific performance, investigate incidents,
and restrict or suspend deployment;
• clinicians can report recurrent errors, unexpected behavior, and clinically significant discrepancies;
• professional societies can develop error taxonomies, clinical performance measures, and specialty-specific
evaluation standards;
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 24
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
• standards-setting bodies can establish interoperable monitoring methods and reporting formats;
• upstream model providers can disclose relevant changes, limitations, and safety information; and
• regulators and public-private consortia can facilitate reporting and signal detection across institutions.
These roles should be defined contractually and operationally, with clear data-sharing and escalation
obligations. Shared participation should not permit the developer of the final clinical function to disclaim
responsibility or transfer accountability to the clinician or institution.
Question 22: How should re-benchmarking scale to the nature of a modification?
The extent of re-benchmarking should be proportionate to the likelihood that a modification could alter
clinical behavior, workflow containment, or the benefit-risk profile.
Changes limited to non-clinical formatting, administrative interfaces, or other functions shown not to affect
clinical output may be managed within the quality management system.
More substantial re-evaluation should be required for changes to:
• model weights or foundation-model versions;
• prompts;
• retrieval sources;
• guardrails;
• clinical-context inputs;
• post-processing;
• report templates that affect clinical content;
• user-interface presentation;
• intended users;
• intended populations;
• autonomy; or
• downstream communication pathways.
Changes that expand the intended use, reduce specialist oversight, introduce action-taking behavior, or expose
unverified output to clinicians or patients may warrant FDA review before implementation.
PCCPs may be appropriate for bounded and foreseeable categories of change when the modification process,
validation methods, acceptance criteria, and rollback requirements are defined in advance.
Question 23: How should PCCPs address future changes that cannot be fully prespecified?
A PCCP need not identify the exact content of every future model modification. It should, however, define
the permissible categories and boundaries of change.
A suitable PCCP may specify:
• the components that may change;
• the intended purpose and maximum scope of those changes;
• prohibited changes;
• the data and development methods that may be used;
• required re-benchmarking and clinical confirmation;
• subgroup and human-factors evaluation;
• acceptance criteria;
• post-deployment monitoring;
• version control; and
• suspension or rollback procedures.
A PCCP should not function as an open-ended authorization for any future modification.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 25
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Changes that fall outside the predefined categories, materially alter intended use, increase autonomy, change
the recipient of the output, or create a new pathway to clinical reliance should undergo the otherwise
applicable regulatory review.
Question 24: How should manufacturers manage changes made by upstream foundation-model
developers?
Manufacturers using third-party foundation models should maintain sufficient technical and contractual
control to prevent an upstream change from entering clinical production without evaluation.
Appropriate mechanisms include:
• version pinning;
• advance notice of material changes;
• detailed change logs;
• defined support and retirement periods;
• audit rights;
• access to safety-relevant performance information;
• automated detection of model-version changes;
• sandbox testing;
• validation against established benchmarks;
• automatic blocking of unvalidated versions;
• rollback capability; and
• contractual remedies when required information is not provided.
Where appropriate, anticipated upstream changes may be addressed through a PCCP.
If a sponsor cannot obtain version pinning, advance notice, or sufficient safety-relevant information from an
upstream provider, that limitation should be disclosed and considered in the benefit-risk assessment.
Silent or automatic replacement of a clinically deployed foundation model should not be permitted when the
change may affect diagnostic behavior or safety. The downstream developer should retain the ability to
continue using a validated version, suspend the function, or transition only after the updated configuration has
satisfied applicable acceptance criteria.
D. Other Topics
Question 25: Would voluntary Foundation Model MAFs be practical and useful?
Voluntary Foundation Model MAFs could improve regulatory efficiency by allowing an upstream model
developer to provide confidential, standardized information that multiple downstream sponsors may
reference.
Useful content may include:
• model architecture and version;
• training-data provenance and general characteristics;
• intended supported use cases;
• known limitations and failure modes;
• healthcare-relevant evaluation results;
• subgroup performance;
• behavioral constraints and guardrails;
• cybersecurity information;
• audit-log capabilities;
• versioning and retirement policies; and
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 26
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
• commitments to notify downstream developers of material changes.
The program will be useful only if files are maintained, review teams can rely on current information, and
downstream sponsors receive sufficient change notifications to manage their own products.
A Foundation Model MAF should not presumptively authorize or classify the underlying model for all
potential downstream uses. The developer of the final configured function should remain responsible for
demonstrating the safety and performance of that function.
Alternative mechanisms may include contractual disclosure obligations, structured model or system cards,
audit rights, independent attestations, standardized APIs for version and change information, and direct
confidential regulatory access to safety-relevant documentation.
Question 26: What additional oversight is appropriate for agentic GenAI-enabled devices?
Agentic AI presents additional risk because it may plan and execute multiple steps, use external tools, and
take actions with fewer opportunities for human review.
Evaluation should therefore address:
• task planning and sequencing;
• compliance with intended-use boundaries;
• tool-selection and tool-use accuracy;
• recognition of erroneous tool outputs;
• authorization and permission controls;
• human-oversight checkpoints;
• handling of interrupted or failed actions;
• prompt-injection and cybersecurity resistance;
• auditability;
• rollback; and
• postmarket detection of unexpected action sequences.
High-consequence, irreversible, or clinically material actions should generally require affirmative human
confirmation unless the autonomous function has been specifically evaluated and authorized for that use.
A qualifying RSG-PD function is not agentic in the clinical sense addressed by this question. It generates
private textual work product and cannot independently communicate, order, treat, escalate, or otherwise act
on the patient's care.
The term "agentic AI" should not itself determine device status. Some agentic radiology functions operate
only on finalized physician-authored reports or administrative data and perform documentation, routing,
coding, communication, or quality-workflow tasks without independently analyzing images or generating
diagnostic conclusions. Each function must be assessed according to its particular intended use, but tool use,
multi-step planning, or agentic architecture alone should not cause a non-device operational function to be
treated as a medical device.
If agentic functions are later combined with pre-drafting — such as autonomous communication, workflow
reprioritization, or order initiation — each action-taking function should be separately identified and
evaluated according to its intended use and risk.
XVII. Workforce and Public-Health Considerations
Patient safety must remain primary. Regulatory cost or operational convenience cannot justify an unsafe
function. FDA should nonetheless consider the public-health consequences of imposing an autonomousdiagnostic evidentiary burden on technology intended only for specialist-supervised drafting.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 27
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Radiology practices face sustained pressure from rising imaging volumes, uneven specialist availability, and a
workforce pipeline that expands slowly relative to demand. Recent American College of Radiology
publications describe persistent supply-and-demand pressure, increasing imaging volumes, and the need to
redesign work rather than merely redistribute workload.24
Developing and validating radiology VLMs also depends on feedback from practicing radiologists, and that
feedback is difficult to obtain outside clinical practice. Radiologists in active practice do not, as a rule, set
aside clinical work to participate in model development; their expertise is applied at the workstation, one
examination at a time. A qualifying RSG-PD workflow makes that expertise available to development
without asking anything additional of the radiologist. Each time a pre-draft is reviewed against the source
images and edited, replaced, or rejected before the radiologist issues the final report, the comparison between
the generated text and the radiologist-approved report is itself a labeled observation, produced in the ordinary
course of practice and at clinical volume. Patient safety is unchanged, because no report issues without the
radiologist’s independent review and signature. If a supervised pre-drafting tool must instead achieve the
standalone performance expected of an autonomous diagnostician before any controlled clinical use, that
source of real-world feedback is unavailable during the period in which it would matter most. Used under the
controlled, offline, versioned process described in Section XIII, this feedback offers a mechanism by which
supervised functions can improve over time without any unverified output ever reaching the patient-care
pathway.
A proportionate framework can protect patients while allowing shadow testing, representative local
validation, limited supervised deployment, collection of real-world error data, controlled improvement, and
expansion only when safety is demonstrated. This is consistent with FDA's stated goal of a scientifically
rigorous, least-burdensome approach that provides reasonable assurance of safety and effectiveness.25
XVIII. Specific Requested FDA Policy
TMI and RADIN respectfully request that FDA develop guidance substantially consistent with the following
principle:
Proposed policy language
FDA intends to exercise enforcement discretion with respect to premarket authorization requirements for
a generative artificial intelligence software function that analyzes medical images solely to generate
private, non-final report work product for review by the qualified specialist physician independently
interpreting those images, when enforceable technical controls, supported by procedural controls, prevent
the output from being finalized, communicated, acted upon, or otherwise entering the patient-care
pathway before affirmative specialist review.
To qualify, the function should include the following conditions, which correspond to those set out in Section
VIII:
1. mandatory specialist access to the complete source images and relevant clinical information before
report finalization;
2. restriction of output to the interpreting specialist;
3. clear identification as unverified AI-generated work product;
Elizabeth Y. Rula, The Radiologist Shortage: A Workforce Update from HPI, ACR Bulletin (Feb. 5, 2026), https://www.acr.org/ClinicalResources/Publications-and-Research/ACR-Bulletin/2026/radiologist-shortage-work-force-update; Andrew Moriarity & Gregory N. Nicola,
Workforce Economics: How Imaging Practices Are Adapting, ACR Bulletin (Apr. 6, 2026), https://www.acr.org/ClinicalResources/Publications-and-Research/ACR-Bulletin/2026/workforce-economics.
FDA GenAI Discussion Paper, supra note 1, at 1-2 (stating CDRH's goal of a nimble, least-burdensome approach that enables timely access
to safe and effective devices); FD&C Act section 513(a)(2), 21 U.S.C. section 360c(a)(2) (reasonable assurance of safety and effectiveness).
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 28
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
4. clinical non-persistence of unverified output outside restricted audit and quality records;
5. prohibition against automatic finalization or passive acceptance;
6. prohibition against downstream communication before review;
7. prohibition against autonomous clinical action;
8. affirmative physician control of the final report;
9. representative local validation;
10. appropriate human-factors evaluation;
11. ongoing performance, drift, and subgroup monitoring;
12. controlled and auditable model changes;
13. rollback and incident-response capability; and
14. exclusion of autonomous, direct-to-clinician, direct-to-patient, action-directing, and time-critical
uses in which meaningful specialist review is not feasible.
These conditions apply specifically to the RSG-PD function. They would not preclude a separately identified
image-based worklist-prioritization function from operating within the same platform, including through a
shared underlying model, provided that all requirements applicable to the prioritization function are satisfied.
The combined configuration should be evaluated for shared-model dependencies, resource contention, and
workflow interactions that could affect safety, performance, or preservation of the safeguards applicable to
each function.26
FDA should retain and exercise appropriate oversight when any of these safeguards is absent or when AIgenerated information becomes available for clinical reliance before specialist verification. If the Agency
determines that broad premarket enforcement discretion is not appropriate, it should consider a streamlined
low-risk pathway or exemption supported by defined special controls rather than requiring supervised predrafting to satisfy standards designed for autonomous diagnosis.
XIX. Conclusion
Generative AI and VLMs may improve radiology quality, consistency, efficiency, and access, but they also
present legitimate risks. Those risks should be addressed directly through validation, human-factors
engineering, access controls, monitoring, cybersecurity, auditability, and controlled change management.
They should not be addressed by treating every model capable of image analysis as though it autonomously
diagnoses patients. A foundation model's latent capability is not the same as the intended use and risk of the
final configured function. A private pre-draft shown only to the responsible specialist is not equivalent to a
preliminary report delivered to a clinician. A reversible textual suggestion checked against complete source
images is not equivalent to an output the user cannot independently verify. A continuously supervised
function is not equivalent to an autonomous function.
A narrowly bounded premarket enforcement-discretion policy for RSG-PD functions would preserve FDA
jurisdiction and patient protections while allowing responsible specialist-supervised innovation. TMI and
RADIN respectfully request that the Agency preserve these distinctions as it develops its regulatory approach.
26 21 C.F.R. § 892.2080; U.S. Food & Drug Admin., Multiple Function Device Products: Policy and
Considerations, §§ V-VII (July 2020), https://www.fda.gov/media/112671/download.
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 29
DRAFTING
TOTAL MEDICAL IMAGING LLC RADIN LLC
Respectfully submitted,
Alejandro Bugnone, M.D.
Chief Executive Officer, Total Medical Imaging, LLC
Chief Executive Officer, RADIN, LLC
Practicing Radiologist
Submitted on behalf of Total Medical Imaging, LLC and RADIN, LLC
JOINT FDA COMMENT • RADIOLOGIST-SUPERVISED VLM PREFDA-2026-N-7874 • PAGE 30
DRAFTING