VitaSignal
“Oversight is meaningful only when it occurs before the consequential action, the reviewer has the information needed to decide, and the system cannot bypass or pressure the checkpoint.”
What they argued
Supports risk-proportionate TPLC; monitoring 'should not substitute' but uncertainty acceptable if detectable/reversible; approval before irreversible actions; pinning and rollback.
Themes it raises
Across the five cross-cutting questions
High-consequence work: Directs
The comment as filed
VitaSignal appreciates the opportunity to comment on FDA’s discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices. We support a risk-proportionate total product life cycle approach.
1. Risk classification should include practical modifiers
The proposed activity and consequence axes are useful, but systems in the same cell can present different risks. FDA should also consider reversibility, time to correction, error detectability, traceability to data and system versions, safeguard independence, exposure before containment, and stability of the population and workflow.
Directiveness should be judged by observed behavior, not interface labels. Personalization, urgency, repeated recommendations, confidence presentation, defaults, omitted alternatives, and automated tool use can move an apparently informational function toward action direction. Multi-turn evaluation should cover context accumulation or loss, escalation, refusal, interruption, and safe stop.
2. Competency evidence should be claim-specific and traceable
Competency should be supported by structured evidence rather than one benchmark score. Each claim should identify the population, setting, user, workflow, input distribution, reference standard, acceptance criteria, failure rules, denominators, uncertainty, version, and exclusions. Testing should include missing, contradictory, shifted, adversarial, and low-quality inputs, plus repeated-run and equivalent-input testing.
Public benchmark performance should not be presumed to predict clinical performance. Sponsors should explain why a benchmark represents the intended use and where its distribution differs from the operating environment. When a sponsor controls test construction, tuning, scoring, and interpretation, external adjudication, preregistration, blinded evaluation, locked tests, or qualified reproduction may be appropriate.
Synthetic data can support schema testing, fault injection, specified rare cases, and reproducibility. It should not, by itself, establish real-world prevalence, workflow effects, subgroup performance, clinical performance, or benefit-risk. Synthetic and real-data results should remain separate unless they estimate the same quantity under compatible conditions.
3. Lifecycle monitoring should add evidence, not replace it
Postmarket monitoring should not substitute for evidence needed before use. Greater residual uncertainty may be acceptable only when failures are promptly detectable, consequences are limited or reversible, exposure is bounded, and containment is feasible.
A credible lifecycle program should include a versioned baseline; denominator and deployment-context data; performance, safety, data-quality, and process indicators; investigation thresholds; incident linkage; review ownership; and rollback, suspension, or reassessment rules. Reassessment should follow material changes in the model, prompts, retrieval corpus, orchestration, tools, guardrails, interface, intended use, users, population, workflow, source data, or clinical practice. New safety signals, subgroup degradation, unexplained output shifts, repeated tool failures, and lost traceability should also trigger review.
4. Foundation-model and agentic dependencies need explicit controls
Manufacturers using third-party foundation models should maintain a versioned dependency inventory rather than rely on a model name or release label. Controls should support change notice, version identification, incident notification, impact assessment, and rollback or pinning where feasible. Regression tests, canary testing, runtime fingerprints, shadow evaluation, and output-distribution monitoring can help detect change. If a dependency cannot be adequately detected, assessed, or controlled, that uncertainty should narrow permitted use and strengthen safeguards.
Agentic systems require evaluation of output quality and action-sequence safety. Criteria should address tool selection, parameters, authorization boundaries, least privilege, stale or adversarial results, approval before irreversible actions, interruption, timeout, rollback, recovery, prompt injection, logging, partial failure, and refusal when authority or confirmation is missing. Oversight is meaningful only when it occurs before the consequential action, the reviewer has the information needed to decide, and the system cannot bypass or pressure the checkpoint.
Conclusion
FDA’s framework would be strengthened by structured risk modifiers, traceable competency evidence, explicit uncertainty, and versioned lifecycle decision records. Calibration, benchmark capability, and monitoring should not conceal failures in intended-use relevance, discrimination, threshold behavior, or clinically important strata. Null, adverse, conflicting, blocked, and unavailable evidence should remain visible throughout the product life cycle.