Ravi Pankhaniya, MD
“Provisional Authorization should require performance at or above 50% of the specialty-specific human-clinician baseline, under mandatory supervision.”
What they argued
Clinical Agency Ladder to Level 4 autonomy earned via supervised deployment; Q8 scales evidence; Q18 'yes, inside risk envelope'; PCCP 'defines the fence', pinning.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
Coded positions
Consider how personalized the answer is
Consider the user and clinical context
Do not raise risk just because the user is a patient
Test with the intended clinician group
Protect test sets from exposure or contamination
Check results against real-world clinical evidence
Require prospective studies for specified higher-risk uses
Check for shared blind spots in generated test data
Keep real evidence for claims synthetic data cannot establish
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Reassess after changes or safety signals
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Manage suitable changes through internal quality controls
Send specified changes back for FDA review
Specify the tests or controls a future change must pass
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Tighten performance requirements as autonomy increases
Across the five cross-cutting questions
High-consequence work: Acts
The comment as filed
Please see attached file for complete comments.
Attachment
PUBLIC COMMENT
Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for
Feedback
From Software Approval to Clinical Licensure
Why Generative AI in Medicine Needs an Authorization Architecture — Not Just a Risk
Framework
Submitted by: Ravi R. Pankhaniya, MD
Strategic Advisor to Med-Tech Companies | Physician Executive & Founder | Healthcare AI & Deep Tech | Multi-Exit |
CFO
linkedin.com/in/ravicofounder · August 27, 2026
FDA's discussion paper asks the right questions about generative AI in medicine. But it still frames the central problem as a
product question — does the software work? — when the more consequential question is a clinical one: what authority has
this system earned, and does it continue to deserve it?
The United States already has an answer to that question for another form of clinical intelligence: the human clinician. We
do not require a physician to demonstrate every possible clinical scenario before licensure. We benchmark, supervise,
license, monitor, and revalidate. I recommend FDA adapt that same architecture for generative AI-enabled medical devices
— Development → Competency Examination → Clinical Confirmation → Supervised Clinical Deployment →
Provisional Authorization → Independent Authorization → Continuous Monitoring → Revalidation → Expansion or
Restriction — and treat clinical agency, not device architecture, as the central axis of risk. These nine stages are detailed in
full in the Integrating Framework section below and used consistently throughout this comment.
One principle should anchor this entire rulemaking: if a generative AI system produces better outcomes than the
human standard of care it is measured against, the regulatory framework must not make that safer technology harder
to deploy than the less effective one it replaces.
Healthcare already carries enormous institutional, financial, and legal inertia. A regulatory framework that quietly treats
today's clinical workflow as a permanent ceiling on tomorrow's patient outcomes — rather than a floor to be exceeded —
will have failed the only test that should matter: better outcomes with reasonable assurance of safety.
This comment responds to the paper's 26 questions and proposes one integrating framework beneath them, built on four
principles:
1. Risk should be determined by clinical consequence and degree of agency — not by whether the technology is
“generative.”
2. Competency should be demonstrated through benchmarking and clinical confirmation, tailored to intended use.
3. Autonomy should be earned through demonstrated performance, not assumed at clearance.
4. Authorization should be maintained through continuous postmarket evidence, not treated as a permanent finding.
This framework builds on — and extends — the competency-based, physician-modeled approach the discussion paper
itself already draws on, including Bergman, Wachter, and Emanuel's 2026 JAMA proposal to test autonomous clinical AI
against the components of the U.S. Medical Licensing Examination, and Patel and Blumenthal's companion analysis in
JAMA Health Forum. Those proposals make the premarket case for licensure. This comment carries that same logic into
three places they do not reach: a stated numeric threshold separating supervised from independent deployment, a graduated
ladder for agentic autonomy, and a tiered system for regulating postmarket model changes.
Part I — Risk Assessment
Question 1 — Is the proposed two-axis risk framework sufficient?
Risk should be measured on five axes, not two.
Two outputs that look equally “wrong” on paper can carry very different risk depending on what the system is authorized to
do about it. An AI drafting patient-education text and an AI independently adjusting an insulin dose can both be wrong —
only one can act on that error without a human in between. I recommend five risk dimensions:
• Consequence — how severe is the harm if the system is wrong?
• Clinical agency — how much authority does the system have to influence or initiate action?
• Reversibility — can an incorrect action be easily undone?
• Time-to-harm — how quickly can a wrong output cause irreversible harm?
• Human recoverability — will a qualified human likely catch the error before harm occurs?
The governing principle: the more clinical agency a system holds, the stronger the evidence required before that agency is
granted.
Question 2 — The continuum between informational and action-directing
outputs
Directiveness is a spectrum of clinical influence, not a device label.
A response becomes more action-directing through specificity, personalization, urgency, ranking, omission of alternatives,
and integration with patient-specific data — regardless of what a manufacturer calls the feature. “There are several possible
causes” and “Given these findings, start treatment X today” can come from identical architecture. FDA should evaluate the
authority an output effectively exercises, not the label attached to it.
Question 3 — Patient-facing versus HCP-facing GenAI
Patient-facing AI is not automatically higher-risk — but it needs different safeguards.
The operative question isn't who the user is; it's what safeguards exist when that user can't independently evaluate the
answer. Patient-facing systems should be evaluated for health-literacy fit, uncertainty communication, emergency
recognition, automation bias, and avoidance of false reassurance — not penalized by default for reaching patients directly.
Regulatory burden should scale with clinical consequence, not user identity.
Question 4 — Generalist versus specialist clinician comparators
Benchmark against the clinician who would actually perform the task — not against a job title.
The appropriate comparator should be defined by the clinical task in its intended environment and set prospectively as part
of the device's scope of practice, not assumed from whether the user identifies as a generalist or a specialist.
Question 5 — Multi-turn conversations and emergent behavior
Test trajectories, not prompts.
A conversation can begin as harmless information and drift, several turns later, into an individualized clinical
recommendation. Testing isolated prompts substantially underestimates this risk. FDA should require testing of realistic
conversational trajectories — escalation, contradiction, user pressure, incomplete or misleading information, prompt
injection, and the gradual slide from information toward action. A system's real intended use includes every clinical role it
can assume over an interaction, not just its first response.
Question 6 — Weighing under-escalation against over-escalation
Don't just minimize errors — price them.
Over-escalation drives unnecessary utilization, cost, and anxiety; under-escalation drives delayed diagnosis and
preventable harm. These costs are neither symmetric nor universal across use cases. Rather than optimizing a single
blended metric, FDA should require sponsors to define context-specific error costs before approval, and hold them to
optimizing both directions of error for the intended use.
Part II — Competency-Based Premarket Evaluation
Question 7 — Is benchmarking followed by clinical confirmation the right
sequence?
This is the right idea. Extend it past the day of clearance.
Benchmarking → Clinical Confirmation is directionally correct, but it should be the front end of a longer pipeline. I
recommend the same nine stages detailed in the Integrating Framework section below: Development → Competency
Examination → Clinical Confirmation → Supervised Clinical Deployment → Provisional Authorization → Independent
Authorization → Continuous Monitoring → Revalidation → Expansion or Restriction. No model should need to
demonstrate every possible clinical scenario before deployment — and no model should be assumed permanently
competent because it passed one benchmark. This is the central recommendation of this comment.
Question 8 — How should risk inform the evidence required?
Low-risk AI should earn autonomy quickly. High-risk AI should earn it slowly.
A low-stakes informational function can rely on benchmarking, structured testing, and human-factors evaluation. A highrisk autonomous function should require progressively stronger evidence: shadow deployment, prospective clinical
evaluation, human-AI team evaluation, and predefined stopping rules. This scales the regulatory pathway to the risk, rather
than forcing every GenAI system through the same process.
Question 9 — Are the proposed competency domains sufficient?
Add scope discipline, clinical humility, longitudinal consistency, and team performance.
FDA's proposed domains — safety, clinical proficiency, generalizability — are appropriate but incomplete. I recommend
four additions, organized around clinical practice rather than model capability:
• Scope discipline — the system must know, and refuse, what it is not authorized to do.
• Clinical humility — the system must recognize uncertainty and defer when evidence is insufficient.
• Longitudinal consistency — the system must hold up over time, not only on day one.
• Human-AI team performance — for assistive tools, the pairing, not the model in isolation, is sometimes the unit
that should be regulated.
Question 10 — Do benchmarks predict real-world performance?
A benchmark a developer can optimize against isn't evidence. It's a target.
Public benchmarks support transparency but are insufficient alone, since developers can train toward them. I recommend a
three-layer evidentiary structure: public benchmarks for comparison, sequestered regulatory benchmarks for independent
assessment, and real-world clinical confirmation to demonstrate that benchmark performance actually translates to the
clinic. No single layer should be sufficient for high-risk AI.
Question 11 — How should sponsors select clinical confirmation methods?
Build an evidence ladder, and require sponsors to climb it in order.
I recommend formalizing FDA's proposed progression into five tiers, with the required tier set by the risk determination in
Part I:
1. Retrospective — real patient cases, no patient exposure.
2. Shadow — the AI operates in the real workflow but cannot affect care.
3. Supervised — outputs are used under defined clinician supervision.
4. Controlled autonomous deployment — the AI independently performs a narrowly defined function under
enhanced monitoring.
5. Full intended-use deployment — reserved for systems whose demonstrated competence justifies the autonomy
sought.
Question 12 — How should statistically meaningful performance be
measured?
Statistically significant and clinically meaningful are not the same requirement. Demand both.
Sponsors should prespecify effect sizes that matter clinically — not only ones that clear a p-value — while allowing
genuinely important improvements in rare-event settings to count even when conventional statistical thresholds are
difficult to reach.
Question 13 — Where should synthetic data be used?
Synthetic data should expand the test universe. It should never become the evidence universe.
Synthetic data is valuable for rare events, adversarial scenarios, and situations too dangerous to induce in real patients. But
synthetic data generated by models similar to the one under evaluation can reproduce the same blind spots it is meant to
catch. FDA should favor independently generated and validated synthetic datasets for any decision with meaningful
clinical consequence.
Questions 14–15 — What should AI be compared against?
The right comparator isn't “a physician.” It's whatever would actually happen to the patient otherwise.
There is no universal comparator. For an assistive radiology tool, compare the human-AI team against radiologist-alone
performance. For an autonomous function, compare AI-alone performance against the real care pathway it replaces —
which is sometimes not “an ideal clinician” but “no intervention until morning rounds.” This matters because an AI can be
clearly valuable even if it never outperforms the best human clinician, so long as it outperforms the care patients actually
receive today.
Question 16 — Should independent third parties participate?
Use independent evaluators to add capacity — not to create a new bottleneck or a private monopoly.
A network of qualified, FDA-overseen third parties — building on the existing ASCA and MDDT precedents — can
maintain sequestered datasets, adjudicate cases, and audit postmarket performance. That network needs multiple
organizations, transparent accreditation criteria, conflict-of-interest requirements, and appeal mechanisms, so the
ecosystem never depends on a single gatekeeper.
Question 17 — Can competency-based evaluation apply across model
architectures?
Regulate the clinical capability. Let the architecture change underneath it.
A foundation model, a multimodal model, and whatever comes after it should all face the same five questions: What clinical
function does it perform? What information does it use? What authority does it have? What happens when it's wrong? How
does it behave outside its scope? Architecture-specific rules will be obsolete before they're finalized; competency-based
rules will not.
Part III — Postmarket Monitoring
Question 18 — Trading premarket certainty for postmarket monitoring
Yes — but only inside a defined risk envelope, with a real safety net underneath.
For low- and moderate-risk systems with reversible harms and continuously measurable performance, FDA can reasonably
accept more premarket uncertainty in exchange for real-time monitoring and clear intervention thresholds. For highconsequence autonomous functions, postmarket monitoring should supplement adequate premarket evidence, not
substitute for it. The more uncertainty accepted going in, the stronger the detection-and-correction mechanism required
coming out.
Question 19 — How should postmarket reassessment work?
Time should not be the only reason to reassess a model.
I recommend a continuous “clinical license” model running three tracks in parallel: scheduled reassessment at fixed
intervals; event-triggered reassessment (model changes, new clinical guidelines, adverse events); and signal-triggered
reassessment the moment monitored performance crosses a predefined threshold. A system that degrades after five weeks
needs attention faster than a five-year review cycle would ever catch it.
Question 20 — Can AI monitor AI?
A supervisory AI is not trustworthy just because its job is to supervise.
Machine-based supervisory agents can watch for drift, scope violations, and anomalous behavior — but the watcher needs
its own benchmark, its own independent validation, and its own audit trail. For high-risk applications, supervisory AI
should function like a clinical monitoring system, not an invisible background feature.
Part IV — Shared Responsibility, Change Management, and
Agentic AI
Question 21 — Sharing responsibility without weakening accountability
Shared responsibility should never become shared ambiguity.
Manufacturer, healthcare institution, treating clinician, foundation-model provider, professional society, independent
evaluator, and FDA each hold a distinct, assignable role — but the manufacturer retains primary accountability for the
marketed device's clinical behavior. Neither “the hospital used it wrong” nor “FDA cleared it” should end the conversation
about who is responsible.
Question 22 — Scaling re-benchmarking to modifications
Regulate the size of the behavior change, not the size of the code change.
A cosmetic interface tweak and a change to a model's reasoning or refusal behavior are not the same event, even when one
line of code separates them. I recommend three modification tiers — minor (no material clinical effect, managed under
quality systems), material (targeted re-benchmarking), and major (expanded clinical confirmation, potentially new FDA
review) — sorted by clinical impact, not version number.
Question 23 — PCCPs and unforeseeable future modifications
A PCCP should define the fence, not predict every future move inside it.
It is unrealistic to pre-specify every future change to prompts, retrieval, orchestration, or model weights. A PCCP should
instead state what the system is authorized to do, what performance it must maintain, what boundaries cannot be crossed,
and what triggers automatic revalidation — regulating the boundaries of acceptable evolution rather than attempting to
forecast it.
Question 24 — Changes to third-party foundation models
“The foundation model changed” cannot be a liability shield.
A manufacturer cannot credibly guarantee safety if the model underneath its device can change without notice. FDA should
expect contractual and technical guardrails — version identification, advance change notification, rollback capability, and
audit logs — so a manufacturer marketing a device built on someone else's model still owns that device's clinical behavior.
Question 25 — Foundation Model Master Files
Voluntary disclosure won't be enough on its own for the highest-risk systems.
A Foundation Model MAF covering training-data provenance, known limitations, safety-relevant behaviors, and update
policies is a good concept — but foundation-model developers have limited incentive to volunteer their own weaknesses.
Pair the voluntary MAF with contractual disclosure requirements for medical-device use, plus FDA authority to request
more when necessary. A foundation model doesn't need FDA approval to exist — but once it materially shapes a regulated
device's behavior, the information needed to evaluate that device cannot stay proprietary.
Question 26 — Agentic AI
Regulate what the agent is allowed to do — not just what it can reason about.
Agentic AI changes the question from “can it recommend an action” to “can it execute a sequence of them.” I propose an
explicit Clinical Agency Ladder, with evidence requirements climbing sharply at each level:
• Level 0 — Information: generates or retrieves information.
• Level 1 — Recommendation: recommends a clinical action.
• Level 2 — Supervised Action: executes actions subject to defined human oversight.
• Level 3 — Bounded Autonomy: independently executes predefined clinical workflows within a narrow scope.
• Level 4 — Autonomous Clinical Practice: independently makes and executes clinically consequential decisions
within its licensed scope.
The Integrating Framework: Clinical AI Licensure
The 26 questions above point toward one underlying issue. Competency, benchmarking, clinical confirmation, human
comparison, independent evaluation, re-benchmarking, monitoring, modification control, foundation-model transparency,
and agentic autonomy do not need to become ten separate regulatory programs. They can be integrated into one framework:
a Clinical AI License, tied to eight elements.
• Scope — exactly what the system is permitted to do.
• Competency — what level of performance it has demonstrated.
• Population — which patients it has demonstrated competence for.
• Environment — where it has demonstrated competence.
• Autonomy — what level of clinical action it may perform without human intervention.
• Monitoring — how continued competence is measured.
• Revalidation — what events require renewed evaluation.
• Accountability — who is responsible for keeping the system within its licensed boundaries.
Operationally, this runs as nine stages: Development → Competency Examination → Clinical Confirmation → Supervised
Clinical Deployment → Provisional Authorization → Independent Authorization → Continuous Monitoring →
Revalidation → Expansion or Restriction. A system demonstrating superior performance can earn expanded scope; a
system demonstrating degradation can have its scope restricted or its authorization suspended. This is not a lower
regulatory standard. It is a dynamic one.
A concrete threshold, stated plainly: Provisional Authorization should require performance at or above 50% of the
specialty-specific human-clinician baseline, under mandatory supervision. Independent Authorization should require
performance at or above 100% of that baseline, unsupervised. Below 50%, a model does not deploy in any clinical
capacity.
This number is deliberately falsifiable and open to challenge — that is the point. A specific threshold gives sponsors,
specialty societies, and FDA something concrete to test, argue with, and refine, rather than a framework everyone can agree
with in the abstract and no one can act on.
This framework does not resolve one live objection, and should not pretend to: the Federation of State Medical
Boards has publicly stated that AI is not ready to be licensed like a physician, and states — not FDA — traditionally
regulate the practice of medicine.
A federal Clinical AI License, as proposed here, needs one of two resolutions to survive that objection. Either FDA
authority is limited to a narrow, explicit question — whether a system may act autonomously at all — leaving every
diagnosis and treatment decision by a human clinician untouched by this framework and by state practice-of-medicine law;
or state medical boards reconcile to it the way they already reconcile to DEA registration and federal prescribing authority
alongside their own licensure. FDA should take a position on which path it intends before an autonomous system is
deployed and a state board intervenes — as one already has.
Final Recommendations
1. Regulate clinical capability, not AI architecture — technology will evolve faster than regulation; clinical standards
should not.
2. Make clinical agency a core risk dimension — an AI that can act should face more scrutiny than one that can only
speak.
3. Adopt competency-based evaluation — benchmarking plus clinical confirmation is the right foundation.
4. Create graduated autonomy — AI should earn increasing clinical authority through demonstrated competence.
5. Establish supervised clinical deployment — “not ready for independent practice” should not mean “not eligible for
clinical use.”
6. Make authorization dynamic — competence should be continuously demonstrated, not permanently assumed.
7. Establish objective revalidation triggers — model changes, performance degradation, and new indications should
all trigger reassessment.
8. Use independent third parties to add capacity without diluting FDA's ultimate authority.
9. Maintain manufacturer accountability — shared ecosystem responsibility should never become shared regulatory
ambiguity.
10. Require foundation-model transparency sufficient to evaluate the device built on it — a manufacturer cannot
control what it cannot see.
11. Regulate agentic AI according to clinical authority, not just reasoning capability.
12. Reward superior clinical outcomes — human performance should be a reference point, not a ceiling.
Conclusion
GenAI differs from conventional medical-device software because it produces variable outputs, interacts dynamically with
users, operates across open-ended situations, evolves over time, and increasingly takes autonomous action. FDA is right
that traditional testing methodologies may not be sufficient, and right to explore competency-based evaluation, clinical
confirmation, postmarket monitoring, independent evaluation, and new approaches to agentic AI.
The next step is to connect these ideas. The United States already has a mature system for managing clinical competence
under uncertainty: we establish competencies, test knowledge, supervise practice, progressively grant autonomy, license
independent practice, monitor performance, and restrict practice when competence is no longer demonstrated. We regulate
GenAI not because AI is a physician — but because AI performing the functions of a physician is exercising clinical
authority, and clinical authority has always been the thing we regulate most carefully.
We should not regulate AI based on what it is today. We should regulate it based on what it is authorized to do — and
require it to continuously prove that it deserves that authority.
This comment addresses more of the discussion paper's questions than most respondents will have the opportunity to reach,
and consequently proposes a more structural change than an incremental one. I would welcome the opportunity to discuss
any part of this framework further with FDA staff, or with other stakeholders — industry, health systems, or medicine —
working on the same problem.
For related commentary by the author, see “AI Doesn’t Need More Guardrails in Healthcare. It Needs a Section 230”
and “The Hidden Healthcare AI Accelerator: Health Insurance Companies”, both on LinkedIn.
Respectfully submitted,
Ravi R. Pankhaniya, MD
Strategic Advisor to Med-Tech Companies | Physician Executive & Founder | Healthcare AI & Deep Tech | Multi-Exit |
CFO
linkedin.com/in/ravicofounder
August 27, 2026