← All 95 filings

VitaSignal

IndustryStartupFiled August 28, 2026591 wordsFDA-2026-N-7874-0048
“Oversight is meaningful only when it occurs before the consequential action, the reviewer has the information needed to decide, and the system cannot bypass or pressure the checkpoint.”

What they argued

RecovryAI’s one-line reading of the filing.

Supports risk-proportionate TPLC; monitoring 'should not substitute' but uncertainty acceptable if detectable/reversible; approval before irreversible actions; pinning and rollback.

Themes it raises

10 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“FDA should also consider reversibility, time to correction, error detectability, traceability to data and system versions, safeguard independence, exposure before containment, and stability of the population and workflow.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“Competency should be supported by structured evidence rather than one benchmark score.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Public benchmark performance should not be presumed to predict clinical performance.”
Proving the device works in real careFDA Q11, Q12, Q13, Q14, Q15
“It should not, by itself, establish real-world prevalence, workflow effects, subgroup performance, clinical performance, or benefit-risk.”
Trading premarket certainty for postmarket monitoringFDA Q18
“Greater residual uncertainty may be acceptable only when failures are promptly detectable, consequences are limited or reversible, exposure is bounded, and containment is feasible.”
Watching the device after it shipsFDA Q19, Q20
“A credible lifecycle program should include a versioned baseline; denominator and deployment-context data; performance, safety, data-quality, and process indicators; investigation thresholds; incident linkage; review ownership; and rollback, suspension, or reassessment rules.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“Manufacturers using third-party foundation models should maintain a versioned dependency inventory rather than rely on a model name or release label.”
Devices that plan and take actionsFDA Q26
“Agentic systems require evaluation of output quality and action-sequence safety.”
Whether human oversight is real oversightFDA Q3, Q4, Q14, Q20, Q21, Q26
“Oversight is meaningful only when it occurs before the consequential action, the reviewer has the information needed to decide, and the system cannot bypass or pressure the checkpoint.”
Security, dependencies and what happens when they failFDA Q1, Q9, Q24
“If a dependency cannot be adequately detected, assessed, or controlled, that uncertainty should narrow permitted use and strengthen safeguards.”

Across the five cross-cutting questions

RecovryAI’s reading of the whole filing. Silence is never counted as opposition.
Patient-facing autonomyShould FDA permit patient-facing AI to act with meaningful autonomy within a defined scope?
Supports with conditions
Proportionate evidenceShould evidence requirements scale with clinical risk rather than a uniform high bar?
Supports
Postmarket relianceCan strong postmarket monitoring justify accepting more premarket uncertainty?
Supports with conditions
Competency evaluationCan a device be evaluated on competency benchmarks and clinical confirmation against clinicians?
Supports with conditions
Change controlCan devices on third-party foundation models be maintained under pre-specified change control?
Supports with conditions
Autonomy acceptedThe highest level this filing accepts
Low-consequence work: Acts
High-consequence work: Directs
Read and coded by RecovryAI readers, September 12, 2026. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

VitaSignal appreciates the opportunity to comment on FDA’s discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices. We support a risk-proportionate total product life cycle approach.

1. Risk classification should include practical modifiers

The proposed activity and consequence axes are useful, but systems in the same cell can present different risks. FDA should also consider reversibility, time to correction, error detectability, traceability to data and system versions, safeguard independence, exposure before containment, and stability of the population and workflow.

Directiveness should be judged by observed behavior, not interface labels. Personalization, urgency, repeated recommendations, confidence presentation, defaults, omitted alternatives, and automated tool use can move an apparently informational function toward action direction. Multi-turn evaluation should cover context accumulation or loss, escalation, refusal, interruption, and safe stop.

2. Competency evidence should be claim-specific and traceable

Competency should be supported by structured evidence rather than one benchmark score. Each claim should identify the population, setting, user, workflow, input distribution, reference standard, acceptance criteria, failure rules, denominators, uncertainty, version, and exclusions. Testing should include missing, contradictory, shifted, adversarial, and low-quality inputs, plus repeated-run and equivalent-input testing.

Public benchmark performance should not be presumed to predict clinical performance. Sponsors should explain why a benchmark represents the intended use and where its distribution differs from the operating environment. When a sponsor controls test construction, tuning, scoring, and interpretation, external adjudication, preregistration, blinded evaluation, locked tests, or qualified reproduction may be appropriate.

Synthetic data can support schema testing, fault injection, specified rare cases, and reproducibility. It should not, by itself, establish real-world prevalence, workflow effects, subgroup performance, clinical performance, or benefit-risk. Synthetic and real-data results should remain separate unless they estimate the same quantity under compatible conditions.

3. Lifecycle monitoring should add evidence, not replace it

Postmarket monitoring should not substitute for evidence needed before use. Greater residual uncertainty may be acceptable only when failures are promptly detectable, consequences are limited or reversible, exposure is bounded, and containment is feasible.

A credible lifecycle program should include a versioned baseline; denominator and deployment-context data; performance, safety, data-quality, and process indicators; investigation thresholds; incident linkage; review ownership; and rollback, suspension, or reassessment rules. Reassessment should follow material changes in the model, prompts, retrieval corpus, orchestration, tools, guardrails, interface, intended use, users, population, workflow, source data, or clinical practice. New safety signals, subgroup degradation, unexplained output shifts, repeated tool failures, and lost traceability should also trigger review.

4. Foundation-model and agentic dependencies need explicit controls

Manufacturers using third-party foundation models should maintain a versioned dependency inventory rather than rely on a model name or release label. Controls should support change notice, version identification, incident notification, impact assessment, and rollback or pinning where feasible. Regression tests, canary testing, runtime fingerprints, shadow evaluation, and output-distribution monitoring can help detect change. If a dependency cannot be adequately detected, assessed, or controlled, that uncertainty should narrow permitted use and strengthen safeguards.

Agentic systems require evaluation of output quality and action-sequence safety. Criteria should address tool selection, parameters, authorization boundaries, least privilege, stale or adversarial results, approval before irreversible actions, interruption, timeout, rollback, recovery, prompt injection, logging, partial failure, and refusal when authority or confirmation is missing. Oversight is meaningful only when it occurs before the consequential action, the reviewer has the information needed to decide, and the system cannot bypass or pressure the checkpoint.

Conclusion

FDA’s framework would be strengthened by structured risk modifiers, traceable competency evidence, explicit uncertainty, and versioned lifecycle decision records. Calibration, benchmark capability, and monitoring should not conceal failures in intended-use relevance, discrimination, threshold behavior, or clinically important strata. Null, adverse, conflicting, blocked, and unavailable evidence should remain visible throughout the product life cycle.