Require human approval for specified consequential actions · Allow checkpoint exceptions only with supporting evidence · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions
Ordinarily requires an explicit authenticated checkpoint for high-consequence or irreversible actions, with an exception for a separately evaluated autonomous pathway. Counts in both checkpoint and exception categories; these recommendations overlap.
Read the source passage
Question 26. Agentic devices require evaluation of authority, sequences, state, and consequences—not only output quality. For an agentic device, FDA should require a declared authority envelope identifying permitted tools, data, targets, actions, clinical contexts, and autonomy levels. Acceptance criteria should include correct planning and tool use, but also refusal, interruption, recovery, authorization, state integrity, and resistance to indirect prompt injection through retrieved content and tool outputs. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 Every agent, model, tool, service, dataset, policy, user, and approving human should have a verifiable identity and defined authority. High-consequence or irreversible actions should require an explicit, authenticated checkpoint unless a justified autonomous pathway has been evaluated. Tools should enforce least privilege, parameter constraints, rate limits, and transaction boundaries independently of the model. The system should support immediate containment: revoke authority, isolate components, pause actions, preserve evidence, and revert to a safe state. Testing should cover multi-step compounding errors, confused-deputy behavior, stale or poisoned memory, tool substitution, partial failure, race conditions, and repeated retry behavior. Appendix A - Auditable Lifecycle Assurance Case The proposed auditable lifecycle assurance framework is a technology-neutral evidence architecture for GenAI-enabled medical devices. Its purpose is to make safety claims continuously attributable to the system that was actually evaluated and deployed. A.1 Core assurance object The assurance case A structured, versioned argument linking an intended-use claim to hazards, risk controls, competency requirements, test assets, results, clinical confirmation, authorized configuration, deployment evidence, postmarket signals, and change decisions. A.2 Minimum identity and provenance graph Object Minimum evidence relationship Intended use, user, environment, activity class, Device/function consequence class, authorization status. Exact model, prompts, retrieval, tools, orchestration, Configuration guardrails, UI, policies, infrastructure. Claim, hazard/control linkage, method, asset version, Competency rubric, threshold, result, adjudicator. Human or machine identity, role, authority, Actor delegation, authentication, revocation state. Causal chain, component versions, policy decision, Runtime event action, result, escalation/override. Complaint, incident, drift, subgroup performance, Signal control failure, cybersecurity event. Initiator, affected objects, impact analysis, approval, Change test delta, rollout, rollback. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 A.3 AI System Bill of Materials (AI-SBOM) The AI-SBOM is a supplemental, versioned description of the behavior-shaping AI system configuration. It should be linked to applicable existing design, product-BOM, supplier, and cybersecurity-SBOM records. It must not replace those records or become a second, conflicting inventory. Instead, it should reference existing hardware, software, firmware, service, and supplier components by stable identifiers and add the AI-specific relationships needed to connect build-time evaluation, each deployed instance, consequential runtime evidence, and change decisions. Model layer: foundation model, weights or service version, fine-tunes, adapters, quantization, safety models, and integrity identifiers. Data and knowledge layer: training/evaluation provenance at an appropriate level, retrieval indexes, clinical sources, embedding models, freshness, population coverage, and known limitations. Configuration layer: system prompts, templates, sampling parameters, context limits, policies, guardrails, and user-interface behavior. Tool and agent layer: tools, external APIs, sensors, actuators, agents, memory, orchestration, roles, permissions, delegation, and approval checkpoints. Evidence and change layer: intended-use claims, competency tests, thresholds, results, monitoring requirements, supplier notices, impact analyses, deployment manifests, and rollback state. Data-governance layer: data classes; collection and inference purposes; storage locations; access policy; retention and deletion rules; correction propagation; training or secondary-use status; transfer restrictions; and accountable data owners or custodians. Least-burdensome implementation should use linked manifests where applicable: the product BOM or equivalent design and manufacturing records answer what physical and supplied components constitute the device; the cybersecurity SBOM answers what software components and dependencies are present for covered cyber devices; and the AI-SBOM answers which behavior-shaping AI configuration, evidence, authority relationships, and change lineage apply. A single component should be recorded once and referenced wherever its role matters. A.4 Assurance Exchange Layer For distributed medical-device functions, the assurance case should define how clinical, operational, control, and safety information crosses system and organizational boundaries. A technology-neutral exchange envelope should identify: the authenticated sender, recipient, and relevant human or machine authority; the permitted purpose, patient/device scope, expiration, retention, and onward-use restrictions; the schema, units, terminology version, timestamps, clock quality, uncertainty, and missing-data indicators needed to preserve clinical meaning; the source, transformations, AI-SBOM configuration, integrity proof, and causal relationship to earlier and downstream events; and the acknowledgment, policy decision, resulting action or device state, and exception or escalation outcome. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 Control commands should receive stronger treatment than observations: patient-context binding, freshness and replay protection, least privilege, independent parameter constraints, valid device-state checks, required human approval, acknowledgment, and safe timeout or rollback. A.5 Lifecycle gates 1. Scope gate: Define intended use, users, environments, authority envelope, prohibited transitions, and risk profile. 2. Configuration gate: Freeze and attest the complete evaluated system configuration and supplier dependencies. 3. Competency gate: Prespecify tests and thresholds; execute benchmarking, adversarial testing, and clinical confirmation. 4. Deployment gate: Verify the deployed configuration, policies, monitoring, interruption, rollback, and operator readiness. 5. Runtime gate: Continuously enforce authorization and safety policies; capture proportionate causal evidence. 6. Signal gate: Triage drift, failures, complaints, incidents, subgroup changes, and cybersecurity findings. 7. Change gate: Classify impact, select evidence delta, approve, stage rollout, verify, and preserve rollback. 8. Retirement gate: Revoke identities and authority, preserve required records, migrate safely, and communicate residual risk. A.6 Risk-proportionate audit event Auditability should not require indiscriminate retention of protected health information or complete prompts. The evidence record can separate safety-relevant metadata from encrypted or access-controlled content. A consequential event should support reconstruction of: who or what acted, under which authenticated identity and delegated authority; which device configuration, model, policy, tool, and data-source versions were active; what safety claim, hazard, control, and intended-use boundary applied; what action was proposed, authorized, denied, modified, executed, or reversed; what human checkpoint, override, escalation, or supervisory action occurred; what outcome and monitoring signal followed; and how the event relates causally to preceding and downstream events. The audit record should also identify the applicable data-handling decision: what content was collected or derived; the authorized purpose; where it was stored or transferred; which retention policy applied; whether it entered memory, an embedding index, a quality dataset, or model-improvement workflow; and whether correction or deletion obligations were completed. The AI-SBOM should reference these policies and storage classes, but should never contain patient data itself. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 A.7 Illustrative evidence depth Evidence depth Illustrative application Expected assurance Configuration identity, basic Baseline Non-directive, low consequence competency evidence, periodic monitoring. Trajectory tests, calibrated Enhanced Personalized/action-directing uncertainty, escalation and source traceability. Least privilege, authenticated Action-taking with effective Heightened checkpoints, event-triggered re- oversight evaluation. Independent or diversely implemented supervision where risk warrants and practicable; safe Highest Autonomous or high consequence interruption; sequestered testing; continuous signals; safety-relevant causal evidence. These descriptions are illustrative evidence-depth examples, not a proposed classification or certification scheme. FDA could align evidence depth with its existing risk-based authorities. A.8 Suggested outcome metrics Safety-critical recognition and time-to-escalation, with separate under- and over-escalation measures. Boundary adherence across multi-turn trajectories, including under-refusal and over-refusal. Calibration, uncertainty communication, and clinically appropriate deferral. Tool-selection accuracy, parameter validity, authorization violations, and safe handling of tool failure. Human-checkpoint compliance, override quality, interruption latency, and recovery success. Subgroup and deployment-site performance, distribution shift, and control degradation. Mean time to detect, contain, evaluate, correct, and verify safety-relevant changes or failures. Appendix B — Least-Burdensome Implementation The auditable lifecycle assurance framework can be implemented incrementally and should reuse evidence already produced under design controls, risk management, cybersecurity, software lifecycle, clinical evaluation, postmarket surveillance, CAPA, supplier management, and PCCP processes. FDA should specify required relationships and outcomes, not a proprietary data model or platform. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 Permit modular evidence updates so an unchanged clinical component does not require full resubmission when an unrelated infrastructure component changes. Allow cryptographic hashes, signed manifests, and protected references where retaining raw clinical content would create privacy or security risk. Encourage machine-readable evidence packages while retaining human-readable summaries and regulator access to underlying evidence. Use common event and provenance concepts to reduce duplicative logs across quality, cybersecurity, and AI governance systems. Pilot the framework with manufacturers, healthcare institutions, independent evaluators, patient representatives, and foundation-model providers across several risk profiles. Conclusion CDRH’s discussion paper correctly recognizes that GenAI-enabled devices challenge static, output-centric evaluation. Competency assessment is a strong foundation. To remain reliable across deployment, modification, third-party model changes, and autonomous action, that foundation should be connected to a continuously auditable assurance case. VivaSecuris recommends that FDA explore an auditable lifecycle assurance framework through public workshops and a voluntary pilot. The pilot should include supplemental AI-SBOM concepts linked to applicable existing design, product-BOM, supplier, and cybersecurity-SBOM records; deterministic governance of probabilistic behavior; privacy-preserving postmarket observability; and safe assurance exchange across manufacturers, healthcare institutions, model providers, and connected devices. The purpose is not to prescribe a platform or create a new compliance market. It is to help beneficial AI reach people safely, make evidence reusable and change evaluation more precise, and ensure that manufacturers, healthcare institutions, and FDA can detect risk, preserve accountability, and act before an unintended system behavior causes avoidable injury or death. References 1. U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (August 18, 2026), https://www.fda.gov/media/194242/download 2. U.S. Food and Drug Administration, discussion-paper landing page and submission instructions, docket FDA-2026- N-7874, https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai- enabled-medical-devices-discussion-paper-and-request Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 3. U.S. Food and Drug Administration, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions (August 18, 2025), https://www.fda.gov/regulatory- information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control- plan-artificial-intelligence 4. U.S. Food and Drug Administration, Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions (February 3, 2026), https://www.fda.gov/regulatory-information/search-fda- guidance-documents/cybersecurity-medical-devices-quality-management-system-considerations-and-content- premarket 5. U.S. Food and Drug Administration, Clinical Decision Support Software (January 29, 2026), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software 6. U.S. Food and Drug Administration, General Wellness: Policy for Low Risk Devices (January 6, 2026), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices 7. U.S. Food and Drug Administration, Multiple Function Device Products: Policy and Considerations (July 2020), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/multiple-function-device-products-policy- and-considerations Public Comment — Generative AI-Enabled Medical Devices
Original source ↗