← All 95 filings

Newton’s Tree

IndustryStartupFiled September 3, 20263,068 words · 1 attachmentFDA-2026-N-7874-0055
“A dashboard without an action process is not a risk control.”

What they argued

RecovryAI’s one-line reading of the filing.

Human approval for high-consequence actions; hazard analysis not two-axis sets evidence; Q18 only when manufacturer can detect and control risk; benchmark is gate not assurance; new foundation model is material change.

Themes it raises

16 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“Reversibility, time pressure, traceability, and available safeguards should modify the risk analysis.”
Whether the user can judge the outputFDA Q3, Q4
“A patient-facing device can have more risk when the patient cannot independently assess the output.”
Escalating too little and too muchFDA Q6
“Failed escalation occurs when the device starts an escalation but the handoff does not work.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“Benchmarking should act as a gate to clinical confirmation. It should not provide the full assurance of safety and effectiveness.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“These examples show why benchmark results alone cannot give sufficient assurance of safety and effectiveness.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“A shadow study is useful when local data, integration, workflow, or latency can change performance.”
Trading premarket certainty for postmarket monitoringFDA Q18
“CDRH can accept more premarket uncertainty only when the manufacturer can detect and control the remaining risk.”
Watching the device after it shipsFDA Q19, Q20
“A dashboard without an action process is not a risk control.”
Who is accountable when something goes wrongFDA Q21
“The manufacturer must remain accountable for the marketed device.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“The manufacturer must know when a third-party foundation model changes.”
Devices that plan and take actionsFDA Q26
“An appointment agent must not change an HR system to create appointment capacity.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“Healthcare professional review can reduce the remaining risk. However, the manufacturer must prove that this control is effective.”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“When the model provider does not permit sufficient control, the manufacturer has an unresolved device risk.”
How this fits rules that already existFDA Q8, Q9, Q16, Q25
“CDRH should continue to use the standard concepts of probability and severity.”
Harm from an output that was not wrongFDA Q1, Q2
“A GenAI device can also fail through omission, harmful agreement, delayed escalation, or an unauthorized action.”
What the rules cost sponsors and the marketNot asked by the FDA
“One organization must not control market access.”

FDA questions it names

Questions this filing names by number.

Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functionsQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices

Coded positions

Where a position was recorded question by question.
Q1Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?
The framework is insufficient
Q2How should the continuum from non-directive to action-directing outputs, and the risk that changes along it, be accounted for?
Consider how personalized the answer is
Q3When clinical information goes straight to the patient, does the risk change, and what safeguards help without underestimating patients?
Base risk on the task and available safeguards
Q4Should it matter whether the clinician using the AI is a generalist or a specialist?
Assess the clinician’s task-specific knowledge
Require specialist review or escalation when needed
Test with the intended clinician group
Enforce user roles rather than relying on labeling
Q5How is risk assessed when a conversation starts with non-directive information and drifts into action-directing?
Test whole conversations, not isolated answers
Enforce limits on what the conversation can do
Q6How should under-escalation be weighed against over-escalation?
Set stricter limits on dangerous missed escalations
Set the trade-off for the clinical context
Test how and when care is escalated
Q11When can a device be confirmed without a prospective clinical study, and what earns that lighter path?
Some uses can be confirmed without a prospective study
Require prospective studies for specified higher-risk uses
Q12How do you get statistically meaningful performance numbers when synthetic inputs are mixed with real ones?
Prespecify how performance and uncertainty are measured
Report synthetic and real results separately
Q13Where is synthetic data good enough, and where is it not?
Use synthetic cases for rare events and stress testing
Check for shared blind spots in generated test data
Keep real evidence for claims synthetic data cannot establish
Q16What role should independent third parties play?
Use independent parties to hold or maintain test assets
Use independent clinical or safety assessors
Control conflicts and keep evaluation open to competition
Q17Does the approach still work for devices built on other model architectures, such as multimodal vision-language models and world models?
Adapt evaluation to the model architecture
Q18Can greater premarket uncertainty about a GenAI device’s benefit-risk profile be accepted through greater reliance on postmarket monitoring?
Allow it only under defined conditions
Q19How should an AI device be monitored after launch, and what sets the cadence?
Track downstream outcomes or execution records
Q20Could AI supervisory agents help carry out postmarket monitoring?
Use AI monitoring with validated safeguards
Q21What roles should clinicians and institutions play in monitoring, without diluting manufacturer accountability?
Keep the manufacturer responsible for investigation and action
Give healthcare institutions a defined monitoring role
Involve societies, standards bodies and other partners
Q24When the foundation model’s developer changes the model, how does the device maker detect it and respond, so safety and effectiveness are not compromised?
Identify and control the model version in use
Detect supplier updates or unexpected behavior changes
Retest changed models or provide rollback
Q25Would voluntary Foundation Model Master Files be practical, and useful in premarket review?
Use them if specified conditions are met
Q26What extra oversight does an AI that plans and acts in multiple steps need?
Require human approval for specified consequential actions
Limit or test what the agent is allowed to do
Evaluate the full sequence of actions and its effects
Keep records that let investigators reconstruct actions

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports with conditions
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
Supports with conditions
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Acts
High-consequence work: Advises
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

Newton’s Tree submits these comments in response to CDRH’s discussion paper on GenAI-enabled medical devices. Our comments come from our work to evaluate, deploy, govern, and monitor clinical AI in healthcare systems. We recommend three separate assurance layers: competency testing, clinical confirmation, and operational assurance. We also include anonymized examples from real deployments. These examples show why benchmark results alone cannot give sufficient assurance of safety and effectiveness.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

FDA Docket No. FDA-2026-N-7874

Response of Newton’s Tree Ltd. to
FDA Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled
Medical Devices

About Newton’s Tree
Newton’s Tree is a global healthcare AI company that enables healthcare providers to
select, test, deploy, and monitor in-house and third-party AI products through its
enterprise AI platform.
The company is led by a senior leadership team with globally unique experience at the
nexus of healthcare, AI technology, and cutting edge research. We work with the
leading health systems that are bending the adoption curve for AI. We do this through
the development and deployment of the very best technology within our vendorneutral ecosystem.

Main recommendation
CDRH should use three separate assurance layers.

Layer 1 Layer 2 Layer 3

Competency Clinical confirmation Operational
The final device must do The device must work in its assurance
its intended task. It must intended population, The device must continue
also remain within its workflow, and technical to work after deployment.
approved scope. environment. The manufacturer must
detect and control changes
in performance.

No layer can replace another.
A benchmark cannot prove clinical effectiveness. A clinical study cannot prove that
performance will remain stable. Postmarket monitoring cannot correct an unsafe
premarket decision.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Risk assessment
Question 1: The two-axis framework
The two-axis framework is not sufficient for risk assessment.
Device activity does not equal the probability of harm. A highly automated device can
have narrow limits and strong controls. An advisory device can cause harm through
automation bias.
CDRH should continue to use the standard concepts of probability and severity.
The risk analysis should include:
The unsafe device behavior.
The possible harm.
The probability of the harm.
The available risk controls.
The remaining risk after the controls operate.
The analysis must include more than an incorrect output. A GenAI device can also fail
through omission, harmful agreement, delayed escalation, or an unauthorized action.

Reversibility, time pressure, traceability, and available safeguards should modify the risk
analysis.
CDRH does not need more graphical axes.

Question 2: General and patient-specific output
CDRH should use one clear division:
General output.
Patient-specific output.
General output does not use information about one patient.
Patient-specific output uses or applies to one patient’s symptoms, history, images,
measurements, records, or conversation.
CDRH should not classify risk by wording, tone, specificity, or degree of recommendation.
These features are difficult to define. Manufacturers can also change the wording without
changing the function.
CDRH should treat patient-specific output as potentially action-directing. The possible
consequence of the output should then determine the risk category.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

A statement that says “consider increasing the dose” can have the same effect as a direct
instruction. A warning or disclaimer does not remove this risk.
Healthcare professional review can reduce the remaining risk. However, the manufacturer
must prove that this control is effective.
The manufacturer must include automation bias in
this evaluation.

Question 3: Patient-facing devices
A patient-facing device can have more risk when the patient cannot independently assess
the output.

However, CDRH should not assume that patients lack useful knowledge or capability. A
well-designed device can improve access and patient control.
The manufacturer should use the following controls:
A clear and limited scope.
Clear sources for clinical information.
Clear information about uncertainty.
Human review before a high-consequence action.
A working clinical escalation process.
Tests for different ages, languages, and levels of health knowledge.
Tests of long and repeated conversations.
Users can also give false information or try to bypass safety controls. The manufacturer
must include this foreseeable use in the risk analysis.
The manufacturer can monitor user and device interactions. The monitoring system can
identify a high-risk conversation and start a human review.
A second AI system can help with this task. However, the manufacturer must also test and
monitor the second system.

Question 4: Generalist and specialist users
Risk increases when safe use needs specialist knowledge and the user does not have
that knowledge.
The intended user must form part of the intended use.
A device for specialist use and a device for generalist use are different devices for
evaluation purposes. The manufacturer must confirm each intended use.
Controls can include:

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Verified user roles.
Role-based access.
Specialist confirmation for specified actions.
Automatic deferral for complex cases.
Monitoring of agreement, rejection, and override by user type.
The device should use verified role information. It should not estimate a user’s specialty
from the conversation.

Question 5: Multi-turn conversations and scope
The manufacturer must assess complete conversations. Tests of single outputs are
not sufficient.
The tests should include:
Long conversations.
Repeated attempts to cross a scope boundary.
Indirect or fictional questions.
Missing or conflicting information.
New conversation sessions.
Stored memory.
Repeated tests of the same clinical case.
The manufacturer should use deterministic filters when the device can measure the
scope boundary.
Current machine-learning devices already use this method. For example, an adult fracture
device can reject pediatric X-rays from DICOM age data.
A GenAI device can use deterministic controls for:
Patient age.
User role.
Input type.
Clinical domain.
Medication class.
Permitted tools.
Permitted actions.
The control should operate outside the generative model where possible.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

When no control stops out-of-scope behavior, the manufacturer must treat that behavior
as foreseeable. The manufacturer must include it in the safety evaluation.

Question 6: Care escalation
The manufacturer must measure three types of error:
Under-escalation.
Over-escalation.
Failed escalation.
Failed escalation occurs when the device starts an escalation but the handoff does not
work.
The device can select the wrong service. The alert can arrive too late. The
responsible person can fail to act.
The manufacturer should measure:
Sensitivity for required escalation.
False-positive and false-negative rates.
Time to escalation.
The selected care destination.
Completion of the handoff.
The clinical and operational result.
A severe under-escalation error needs its own pass or fail limit. Good performance on
common cases must not compensate for a severe failure.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Layer 1: Competency
Questions 7 to 10: Device benchmarking
The competency approach is useful and necessary.
The manufacturer must test the final device configuration. The test must include the
model, prompts, retrieval system, controls, tools, and user interface.
Benchmarking should act as a gate to clinical confirmation. It should not provide the full
assurance of safety and effectiveness.

Each high-consequence failure must have a separate acceptance limit. An average score
can hide an important failure.
Competency tests should include:
Clinical knowledge.
Important omissions.
Scope control.
Refusal and deferral.
Long conversations.
Escalation.
Repeated output variation.
Input data errors.
Tool use.
Unauthorized actions.
Behavior after a system change.
The manufacturer must map each test to the intended use or to a specified hazard.
The manufacturer should also show that:
The test population represents the intended population.
The test includes rare and severe cases.
The acceptance limits were specified before the test.
The test data were separate from development data.
Qualified independent reviewers assessed open outputs.
Repeated tests gave stable safety results.
Clinical confirmation supports the benchmark results.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

The two-axis framework should not directly set the amount of evidence. The hazard
analysis and the effectiveness of the risk controls should set the evidence requirements.

Question 17: Other model types
The three-layer approach can apply to different model types.
A multimodal device needs tests for:
Incorrect links between text, images, and other data.
Missing or conflicting inputs.
Incorrect patient or time-point matching.
Unsupported statements about an image or signal.
Changes in imaging equipment or acquisition method.
A predictive world model needs tests for:
Error across multiple predicted steps.
Incorrect counterfactual results.
Uncertainty in future states.
Failure when real events differ from predicted events.
The final device remains the main unit of evaluation.

Question 25: Foundation Model MAFs
A Foundation Model MAF can provide useful dependency information. It cannot prove that
a final device is safe or effective.
Useful MAF content includes:
The model and version identifier.
The supported data types.
The intended and unsupported uses.
Known failure modes.
Safety controls.
Healthcare test results.
Output variation.
Change history.
Update notice requirements.
Incident information.
Version control and rollback options.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Service withdrawal plans.
The manufacturer must still test the final device.
A voluntary MAF can reduce repeated work. However, CDRH should require another
method when the model provider does not give sufficient information.

Question 26: Agentic devices
Agentic devices need additional tests and controls.
The manufacturer must assess:
Incorrect plans.
Errors that increase across several steps.
Incorrect tool selection.
Incorrect tool inputs.
Unauthorized actions.
Prompt injection.
Privilege escalation.
Failure to stop.
Failure to reverse an action.
The device should use:
Least-privilege access.
Deterministic tool allowlists.
Deterministic action limits.
Human approval for high-consequence actions.
Complete logs.
Rate limits.
A stop control.
A rollback method.
An appointment agent must not change an HR system to create appointment capacity. The
device permissions must match the approved function.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Layer 2: Clinical confirmation
Question 11: Selection of the confirmation method
The confirmation method should depend on the possible harm and the
remaining uncertainty.
A retrospective study can be sufficient for a narrow function with an objective
reference result.
A shadow study is useful when local data, integration, workflow, or latency can
change performance.

A prospective study can be necessary for an irreversible, time-critical, or highconsequence action.
The manufacturer should use more than one site when population, equipment, workflow,
or user behavior can affect performance.
The manufacturer should report natural clinical prevalence. The manufacturer should also
test enough rare and severe cases.

Question 12: Statistical measurement
The manufacturer must define the unit of analysis.
The unit can be an output, a conversation, an encounter, a patient, an action, or a
clinical result.
The analysis must account for repeated patients, users, and sites.
GenAI output can vary for the same input. The manufacturer must repeat tests and
measure this variation.
The manufacturer should report:
Confidence intervals.
Results for important subgroups.
Results by site.
Results by user type.
Results by model and configuration version.
Results for common and severe cases.
The manufacturer should report synthetic and real data separately.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

The manufacturer should combine the results only when it proves that the two data
sources measure the same performance distribution.
Benchmark results and clinical results should usually remain separate.

Question 13: Synthetic data
Synthetic data are useful for:
Scope tests.
Rare but well-defined hazards.
Numerical errors.
Unit errors.
Adversarial tests.
Tool failures.
Controlled changes to one patient feature.
Synthetic data are not a sufficient replacement for real data about:
Clinical communication.
Human behavior.
Long-term reliance.
Local workflow.
Real acquisition errors.
Poorly represented groups.
Clinical outcomes.
A model from the same model family can reproduce the same blind spots. The
manufacturer should use an independent generator where possible.
The manufacturer must confirm important synthetic test results with real data.

Questions 14 and 15: Comparators
The manufacturer should use an objective reference result when one exists.
For open outputs, qualified clinicians should use a specified scoring method. The
method should assess errors, omissions, actions, uncertainty, source support, and
user understanding.
Specialists should set the reference result for specialist tasks.
The standard of care should set the minimum safety level. Median clinician performance
can provide additional comparison data.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

The manufacturer must assess human-AI team performance when the intended use
includes human review. The manufacturer should also report device-only performance.
The comparator should represent the care that would occur without the device. This can
include unaided judgment, delayed specialist review, current software, or no intervention.
Relevant results include safety, time to care, workload, unnecessary treatment, access,
and subgroup differences.

Question 16: Independent third parties
Independent third parties can:
Keep hidden test data.
run benchmark tests.
provide clinical review.
confirm sponsor methods.
perform multi-site studies.
audit monitoring programs.
The third party must have suitable clinical, statistical, technical, and humanfactors knowledge.
The third party must disclose conflicts. Payment must not depend on a positive result.
FDA should permit multiple qualified organizations. One organization must not control
market access.

The manufacturer must remain accountable for the device.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Layer 3: Operational assurance
Question 18: More postmarket reliance
CDRH can accept more premarket uncertainty only when the manufacturer can detect and
control the remaining risk.

The device should have:
A limited initial deployment.
A safe fallback process.
Continuous version records.
Defined monitoring measures.
Defined alert limits.
Named persons who review alerts.
A rapid stop or rollback method.
A plan to collect the missing evidence.
This approach is not suitable when harm is irreversible or time-critical. It is also not
suitable when the manufacturer cannot detect failure before harm.

Question 19: Postmarket monitoring
Postmarket monitoring should include five areas:
System operation.
Input data quality.
Device performance.
Human and device interaction.
Clinical and operational results.
The manufacturer should monitor availability, latency, failed requests, and
software versions.
The manufacturer should monitor missing data, format changes, acquisition changes, and
input distribution changes.
The manufacturer should monitor errors, omissions, refusals, escalation, and
subgroup performance.
The manufacturer should monitor agreement, override, automation bias, and completion of
recommended actions.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

The manufacturer should also monitor downstream outcomes.
The manufacturer should use both random review and risk-based review. Riskbased review should include unusual inputs, disagreements, severe cases, and
long conversations.
A monitoring signal must have a specified action. Possible actions include review,
restriction, pause, rollback, or withdrawal.
A dashboard without an action process is not a risk control.

Question 20: Machine supervisors
A machine supervisor can find unusual behavior and select cases for review.
The manufacturer must test the supervisor as a safety component.
The test should include:
Sensitivity.
False-positive rate.
False-negative rate.
Subgroup performance.
Prompt-injection resistance.
Common failure with the main device.
Behavior when the supervisor is unavailable.
The supervisor should use a different model family where possible.
The manufacturer should use deterministic controls for fixed safety boundaries. A
probabilistic supervisor should not replace these controls.
Human reviewers should assess samples of flagged and unflagged cases.

Question 21: Stakeholder roles
The manufacturer must remain accountable for the marketed device.
The manufacturer must manage monitoring, investigation, corrective action, reporting,
model dependencies, and change control.
The healthcare organization must manage local data, integration, workflow, access,
training, and incident escalation.
Clinicians should report and review possible errors. They should not be the only
monitoring method.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Professional societies can define specialist hazards, clinical measures, and
reviewer qualifications.
Standards bodies can define common formats for logs, versions, incidents, and
performance measures.
FDA can support a confidential multi-site monitoring network. This network can find rare
signals that one manufacturer or hospital cannot find.
The Newton’s Tree disagreement example also supports shared review. Approximately half
of the reviewed outliers came from the human reference or workflow.

Questions 22 and 23: Changes and PCCPs
The manufacturer should classify changes by their possible effect.

Low-effect changes
The quality system can control changes that do not change device behavior. The
manufacturer must document the assessment and test the change.

Approved bounded changes
A PCCP can control specified changes to prompts, retrieval data, or guardrail limits.
The manufacturer must complete the specified tests before deployment. The manufacturer
should use a limited initial release.

Material changes
A new foundation model, population, modality, tool, action, or level of autonomy can be a
material change.
A material change can require full testing, new clinical confirmation, and FDA review.
A PCCP should specify:
The components that can change.
The permitted purpose of each change.
The limits of each change.
The required tests.
The acceptance limits.
The deployment method.
The monitoring method.
The rollback criteria.
A change outside these limits should require a new submission.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Question 24: Third-party model changes
The manufacturer must know when a third-party foundation model changes.
Contracts should require:
Advance notice.
Model and version identification.
Release information.
Safety incident information.
A suitable deprecation period.
Access to an earlier version.
Investigation support.
Rollback support.
Technical controls should include:
Version logs.
Fixed production versions where possible.
Automated regression tests.
Tests before migration.
Detection of unexpected behavior changes.
Limited initial deployment.
Rapid rollback.
The manufacturer must assess the complete device after each important model change.
When the model provider does not permit sufficient control, the manufacturer has an
unresolved device risk.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices
FDA Docket No. FDA-2026-N-7874

Conclusion
CDRH should use three assurance layers.
Competency must show that the final device can perform its task and remain
within scope.
Clinical confirmation must show that the device works in the intended
clinical setting.
Operational assurance must show that the device continues to work
after deployment.

CDRH should use a clear division between general output and patient-specific output. It
should not regulate small differences in wording.
Manufacturers should use deterministic controls for boundaries that a system
can measure.
Postmarket monitoring must connect each signal to an accountable action. The available
actions must include investigation, restriction, pause, rollback, and withdrawal.

Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices