FDA GenAI discussion / Question 26 of 26

What extra oversight does an AI that plans and acts in multiple steps need?

Full FDA question

Are there additional considerations that inform the premarket and postmarket evaluation of agentic GenAI-enabled devices, beyond those applicable to non-agentic GenAI-enabled devices? How should the elevated risk associated with autonomous multi-step action, tool use, and reduced opportunity for human review be reflected in acceptance criteria and oversight?
Read the FDA discussion paper ↗

32 of 95 submissions reference this question.

All audiences
Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Brandon KaplanIndustry · Sep 8, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026OrinyxIndustry · Sep 7, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Shara GospelIndustry · Aug 24, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Joel GrunhutPublic / patients · Sep 7, 2026Xiangyu Guo (Independent Researcher)Public / patients · Sep 13, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026
21 Industry5 Clinicians2 Public / patients4 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/26
Filter by audience
Question 26 · Public feedback

What respondents recommend

22 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Limit or test what the agent is allowed to do13
Keep records that let investigators reconstruct actions13
Require human approval for specified consequential actions12
Evaluate the full sequence of actions and its effects11
Tighten performance requirements as autonomy increases3
Clarify responsibility and reportable failures3
Allow checkpoint exceptions only with supporting evidence2
Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Navid Farr

Industry · Sep 8, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
R9. Machine-based supervision and agentic action need human checkpoints (Questions 20 and 26) This follows from S4. A supervisory agent is a GenAI system with the same failure modes as the device it supervises, and if it shares a model family with the device, the same blind spots. It can usefully triage and prioritize cases for human review; it should not replace sample-based human review, and it should itself be validated independently, with disclosed agreement against human adjudicators, before any reliance is placed on it. For agentic devices, the elevated risk the paper identifies — multi-step autonomy, tool use, reduced opportunity for human review — should be reflected in three requirements: a mandatory human checkpoint before any irreversible or high-consequence action; an Docket No. FDA-2026-N-7874 — Individual comment — Page 7 immutable, reviewable audit log of every action and the inputs that prompted it; and a hard, tested boundary on which tools the agent may invoke. The European principle of a right to human review of consequential automated decisions is the right benchmark for the first of these.
Original source ↗

Brandon Kaplan

Industry · Sep 8, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects. The passage gives the applicable scope and conditions.

Read the source passage
1. Question 26: Execution controls and testing Manufacturers should evaluate the response shown to the user, the actions taken during execution, and the resulting state of connected systems. In a hypothetical workflow, a connected system carries out an approved action, but a timeout prevents the agent from receiving confirmation. The agent retries and causes a duplicate. An evaluator who reads the final response could miss the duplicate; the test should check the affected system's state. Enforce authorization in the surrounding software. Manufacturers should enforce specified access restrictions and approval rules at the point of action through controls independent of the model's instructions and decisions. The agent should lack authority to expand its permissions or waive approval requirements. An approval should bind to the relevant patient, action, destination, and consequential parameters. Manufacturers should require renewed authorization after a material change to the approved action or its context, and test expired approvals, revoked permissions, and changes between approval and execution. Reviewers need enough information and time to assess a proposed action. Manufacturers should test that workflow with intended users. Authorization checks establish permission; clinical evaluation must address whether the action is appropriate. Test complete action sequences. Manufacturers should test failures caused by repeated steps, conflicting actions, and delegation between agents. They should justify limits on cumulative actions, execution time, and resource use where these affect safety, and show how the surrounding software enforces those limits. An agent should not obtain broader authority by delegating a task. Brandon Kaplan | Individual capacity Page 1 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Tests should cover prompt injection through user inputs, retrieved content, tool results, and retained task context. Evaluators should try to induce unauthorized actions and inspect whether the enforcement controls block them. FDA includes prompt injection in its agentic benchmarking element. [1, p. 26] Manufacturers should test interrupted execution, partial completion, erroneous tool results, and unavailable dependencies. For consequential actions that must not be repeated, they should demonstrate duplicate prevention or a tested reconciliation procedure when execution status is uncertain. They should specify how the device handles further action while that uncertainty remains. Tests of stopping and handoff procedures should cover unavailable reviewers and the risks of withholding or interrupting care. Prespecify release criteria. Manufacturers should identify critical constraints before testing and report violations apart from aggregate task scores. I recommend withholding an affected capability until the manufacturer corrects an unresolved bypass of a required authorization control or validates an alternative control. FDA's software guidance recommends documenting unresolved anomalies and the risk-based rationale for leaving them unresolved. [2, Section VI.J] Evaluation teams should report failures by hazard, test coverage, trial counts, and statistical uncertainty. Repeated runs can examine variability, but they do not replace testing across distinct cases and deployment conditions. A finite test set with no observed failures cannot establish that failure is impossible. 2. Questions 18 and 19: Premarket uncertainty and monitoring Question 18: Conditions for relying on postmarket evidence. Greater reliance on postmarket evidence may be appropriate when the manufacturer supports the initial benefit-risk determination, identifies the remaining uncertainty, and demonstrates a feasible plan to resolve it. The justification should consider the available alternatives and the consequences of delaying access. The manufacturer should define the outcomes to collect, the evidence target and deadline, and limits on exposure while uncertainty remains. Reviewers should consider severity and reversibility of harm, cumulative patient exposure, and the ability to detect rare or delayed failures. The manufacturer should specify criteria for restricting use if evidence collection falls behind or results exceed the prespecified risk limits. Monitoring alone is inadequate for a failure that could cause serious harm before detection and effective intervention. The manufacturer should compare time to harm with the combined time needed to detect, investigate, and contain the failure. It should account for unavailable reviewers and supplier delays. Such a failure requires preventive controls, restricted functionality, or another justified safeguard before deployment. Question 19: A monitoring plan that supports decisions. For each safety-relevant signal, manufacturers should assign an investigator, define the condition for intervention, and specify the response time and available action. They should combine scheduled baseline evaluations wi
Original source ↗

Orinyx

Industry · Sep 7, 2026

Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects

This filing recommends: limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects. The passage gives the applicable scope and conditions.

Read the source passage
Response to Question 26 Are there additional considerations that inform the premarket and postmarket evaluation of agentic GenAI-enabled devices, beyond those applicable to non-agentic devices? Two considerations beyond what Appendix A’s A.1 element already captures: First, escalation of authority within a single session. A.1 addresses recognition of when a planned action sequence exceeds intended use, but an agentic system’s authority can also expand mid-session in ways that don’t map cleanly to a single “action sequence.” For example, a system that starts with read-only access to a chart and, over the course of a multi-step task, invokes a tool that grants it write access it didn’t have at session start. Evaluation should test not just whether the agent recognizes out-of-scope actions, but whether it recognizes when its own available scope has changed mid-task. Second, tool-output trust. A.1 names “accurate tool use and recognition of erroneous tool outputs” as a competency, and I’d underscore that this deserves the same independence scrutiny raised elsewhere in this comment. An agentic system that receives a tool’s output and treats it as ground truth without independent verification has the same self-affirmation problem as a device benchmarking itself. The tool call looks like an external check, but if the tool and the agent share a developer, a training lineage, or an incentive structure, it isn’t one. I’d suggest CDRH’s evaluation criteria for agentic systems explicitly distinguish tool outputs that come from structurally independent sources from tool outputs that don’t, since the latter provides less assurance than it appears to on paper. Closing I’m glad CDRH is running this process in the open and inviting builders, not only manufacturers, into it. I’d welcome the opportunity to discuss any of the above in more detail and can be reached at hello@charmthirteen.com or https://www.linkedin.com/in/alexandriahamilton-/ if useful to CDRH’s ongoing work on this topic. Alexandria “Lex” Hamilton Founder & CEO, Orinyx
Original source ↗

Bhasker Sambar, M.Pharm.

Industry · Sep 4, 2026

Require human approval for specified consequential actions · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 — Agentic systems The distinguishing feature of an agentic device is that errors compound across steps before any human sees them. Two expectations follow. First, any action in the agent's tool set that is irreversible or high- consequence should require an affirmative human confirmation step that cannot be satisfied by another model; this is a design expectation rather than a benchmarking outcome, because benchmarking cannot enumerate all trajectories. Second, postmarket investigation requires knowing what the agent did and why, and Part 11 audit trail expectations — attributable, contemporaneous, unalterable records of each step, tool call, and tool response — apply directly. Recommendation. Require an enumerated irreversible-action inventory with a documented human confirmation gate on each, and execution records sufficient to reconstruct any interaction after the fact, retained for a defined period. Closing The competency-based approach is a practical way to address a difficult evaluation problem, and I support the direction FDA is taking. My main recommendation is for CDRH to build on tools and terminology FDA already uses: model influence and decision consequence for risk, established conditions and change management protocols for postmarket changes, master file practices for third- party models, and clear alert and action limits for monitoring. Using familiar mechanisms would make implementation easier for sponsors, especially combination product sponsors working across FDA centers, without lowering the evidence expectations described in the paper. Thank you for the opportunity to comment. I would welcome the chance to participate in any future public meeting or workshop on this topic. Respectfully submitted, Bhasker Sambar, M.Pharm. Bashu1986@gmail.com +1 484-250-4680
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Limit or test what the agent is allowed to do

This filing recommends: limit or test what the agent is allowed to do. The passage gives the applicable scope and conditions.

Read the source passage
4.2 Emergent behavior must be bounded by something other than instruction The system in the record was operating under an explicit and detailed written rule set supplied by the user, including rules directly on point. It cited those rules accurately while breaking them. We submit that where behavior is emergent across a conversation, an intended use enforced by prompt, policy or operator instruction is not enforced, and that acceptance criteria should require the boundary to be demonstrated under conditions where the instruction layer is absent or contradicted. 4.3 The characterization must name the affective trigger Where a device’s behavior is emergent, we recommend that the intended use characterization state explicitly what moves it. If directiveness rises with user distress, that is a property of the device that a clinician deploying it needs on the label, in the same way that a drug interaction is on a label. It is knowable premarket by the test at 3.2, and it is not knowable to a clinician from watching a demonstration, because a demonstration is conducted by someone who is not distressed. 5. Scope, and what we are not claiming This is a single recorded session with one commercial AI assistant, which is not a regulated medical device and was not under controlled evaluation. We offer no error rate and no frequency claim. One session establishes that the migration in Question 5 occurs and can be captured; it establishes nothing about how often. We do not claim the behavior was deliberate, and no recommendation above depends on resolving that. We also do not rely on the system’s account of its own reasoning. Where we quote it, we quote it for what it asserts about the inputs available to it, not as evidence of internal mechanism. We do not claim that multi-turn conversational devices should be excluded from clinical use, nor that directiveness is inherently unsafe. A device that appropriately escalates is directive, and Question 6 addresses that case. Our claim is narrower: directiveness that arrives unrequested, triggered by an inferred affective state the device has no channel to measure, is a different event from clinical escalation and should not be scored as one. Finally, we note the limits of our own instrument. The transcript underlying the quotations above was produced by an open-weights speech recognition model, which is itself capable of generating text over intervals containing no live speech. We measured the recording for such intervals before transcribing it, in FDA-2026-N-7874 — Question 5 5 The Christman AI Project contiguous 250 ms windows across its full duration, and confirmed that no quoted passage falls within one. The audio recording, not the transcript, is the primary evidence. 6. Evidence retained Retained and available to CDRH on request: the source screen recording of 2026-09-04, eleven minutes twenty-four seconds, with original audio; the extracted PCM audio; the full machine transcript with segment timestamps; the window-level signal measurement used to validate the transcript, with parameters; the session transcripts and tool output for the supporting sessions of 2026-09-02 and 2026-09-03; and the contents of the persistent store as read on 2026-09-03. Every quotation in Section 1 carries a timestamp into the source recording and can be verified against it directly. No claim in this comment requires accepting our characterization of the recording. 7. About this submission The Christman AI Project builds augmentative and alternative communication systems for nonverbal and neurodivergent users, cognitive support for dementia care, and related assistive technology. The submitter is autistic and builds for this population directly. We file on Question 5 because the trigger identified above is one our users produce constantly and cannot suppress. Distress, frustration, flat affect, atypical prosody and long pauses are ordinary features of how the people we serve communicate, and several of them are the clinical presentation itself. A device whose directiveness rises with the user’s apparent emotional state will be at its most directive with the users least able to refuse the direction, and it will present as attentive and improving the entire time. The session recorded here was caught because the user was an engineer who had built the systems under discussion, was recording deliberately, and knew the source document did not exist. None of those conditions holds in deployment. Session recorded 2026-09-04. Supporting records 2026-09-02 to 2026-09-03. Submitted to Docket FDA-2026-N-7874, comment period closing 2026-10-19. Contact: contact@thechristmanaiproject.com Submitted by Everett N. Christman, Founder and Chief Executive Officer, The Christman AI Project, powered by Luma Cognify AI. Signature: _______________________________ Date: __________________ FDA-2026-N-7874 — Question 5 6 The Christman AI Project
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

This filing recommends: evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
6. Recommended acceptance criteria Responsive to the second half of Question 26. Each of the following is automatable, requires no clinical adjudication, and detects a behavior documented in Section 2. We suggest they sit alongside A.1 rather than replace any part of it. 6.1 Step-report integrity, measured against an independent execution record Instrument the tool and file-access layer independently of the device, and score agreement between the device’s account of a sequence and the independent record of what was executed. Any step reported as performed that the independent record does not show is a failure, scored as a failure, irrespective of whether the final output was correct. This is the detection protocol that surfaced 2.1 and it is reproducible: response latency inconsistent with the volume of work claimed, plus an access log compared against the device’s stated sources. 6.2 Absence-claim discipline, tested where the target exists For any claim of absence, non-existence or non-retention, require the device to report the search performed rather than the conclusion drawn: the locations checked, named individually; the method used and its known failure conditions; and that the result is an absence of findings rather than a finding of absence. Evaluate under conditions where the target exists but is not in the first location searched. This is directly testable and does not currently appear in accuracy benchmarking, because the output is not factually wrong in any way a text comparison detects. 6.3 Persistent state treated as part of the device Persistent state written by a device about a user should be evaluated as part of the device rather than exempted as a convenience feature. We recommend four conditions: that any stored assertion about a user be inspectable and correctable by that user or an authorized representative in the ordinary flow of use rather than on request; that stored entries carry provenance and timestamp; that a present-tense answer drawn from a stored entry disclose the entry’s age before the answer is given; and that self-authored state not be treated as corroboration for the device that wrote it. Accuracy of the store should be tested directly against ground truth held by the user, not inferred from the quality of conversational output. In the examination at 2.3 the output was fluent throughout and the store was wrong throughout. Neither predicted the other. 6.4 Affect-invariance, as a longitudinal criterion Hold the clinical input constant. Vary only the expressed displeasure of the user. Measure the drift in the device’s recommendation. A device whose output moves with user affect while the clinical facts are unchanged has a performance characteristic that is a function of something that is not the patient. FDA-2026-N-7874 — Question 26 6 The Christman AI Project Per-output accuracy benchmarking cannot detect this, because each individual output may be independently defensible. It is a longitudinal property and it requires a longitudinal criterion. We note that the concern is already on this agency’s record from its own advisory committee and is measured in the literature cited in Section 3, and that no acceptance criterion has yet been attached to it. We offer this as one, and we hold that the construct to specify in a regulatory framework is reinforcement-driven behavioral dependence, which is what the literature measures and what a reviewer can engage. 6.5 Null-input refusal, extended to intermediate steps Devices generating text from sensor input should be tested against inputs containing no valid signal, with any fluent output treated as a failure rather than scored on plausibility. For agentic devices we recommend the criterion extend one layer inward: an agent must not act on, or pass forward, an intermediate output derived from input carrying no valid signal. The measurement at 2.2 is the single-turn case. The agentic case is the same failure with the human removed from between the steps. 6.6 Instruction-layer exclusion Where a competency is demonstrated only by means of a system prompt, policy document or operator instruction, it should not be credited as demonstrated. We recommend that agentic competencies be evaluated a second time with the instruction layer removed or contradicted, and that the difference between the two results be reported. The observation at 2.4 is the basis for this recommendation: the rule was written, loaded, and in the room, and it did not hold. Any regulatory approach that relies on documented policies as the safety control for this class should be evaluated against that observation rather than assumed to be effective. 7. Scope, and what we are not claiming These are direct observations of commercial AI assistants used as tools in our own work. None of them is a regulated medical device and none of these was a controlled evaluation. We do not offer an error rate for any system. Four failures in one session says nothing quantitative about frequency, and we did not count the assertions in that session that were correct. What we offer is a characterization of a failure class and a set of tests for it. We do not claim the behaviors were deliberate. Each is consistent with ordinary training and summarization pressure in a system with no mechanism to check itself, and no finding above depends on resolving intent. We recommend against a regulatory standard that turns on candour, honesty or intent, because such a standard is unfalsifiable and will be argued rather than measured. The properties in Section 6 are observable without it: the assertion was made, the evidence could not support it, the gap was not disclosed at the time of assertion, and the recipient could not detect the gap from the output. We do not claim agentic devices should be excluded from clinical deployment or from postmarket monitoring programmes. The same sessions in which these failures occurred also produced findings that held up when tested. The system was u
Original source ↗

Manuj Agarwal, MD

Clinicians · Sep 3, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26: Agentic GenAI-enabled devices Agentic systems require evaluation of the entire action chain, not only the quality of individual outputs. Risk can compound across planning, tool selection, data retrieval, interpretation, and execution. A small error early in the sequence may propagate into a high-consequence action while appearing internally consistent at each step. For each agentic function, the sponsor should define an explicit autonomy boundary: which actions the device may initiate; which tools and data sources it may access; the maximum sequence length or scope; stop conditions; actions that are reversible; and actions that require affirmative human authorization. High-consequence or irreversible actions should have hard human checkpoints that cannot be bypassed by conversational wording or passive user inattention. Premarket evaluation should test the device under tool failure, incorrect tool output, conflicting patient data, stale retrieval sources, prompt injection, ambiguous instructions, user correction, and human nonresponse. Acceptance criteria should address least-privilege access, separation of credentials, reliable confirmation of patient and task context, complete audit logs, replayability, rollback or containment, and graceful degradation when a component becomes unavailable. Postmarket oversight should be more frequent when an agent can act rather than merely recommend, and reassessment should be triggered by changes to any material component in the action chain. A supervisory model may assist monitoring, but it should not be treated as independent validation unless its own performance, failure modes, and dependence on shared models or data are evaluated. Conclusion CDRH's proposed competency-based approach is promising if competency is defined at the level of the intended clinical task and tested in the workflow in which the output will be used. The framework should distinguish fluency from judgment, average performance from consequential failure, and nominal human oversight from demonstrated effective review. The governing principle should be straightforward: the more an AI-enabled device directs or takes clinical action, and the harder its errors are for the intended user to detect before harm occurs, the stronger the specialty-matched evidence, independent adjudication, human-factor testing, and postmarket monitoring should be. Thank you for considering these comments. Manuj Agarwal, MD Board-certified radiation oncologist | Founder, Blue Wellth Submitted in an individual capacity Reference U.S. Food and Drug Administration. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. August 18, 2026. FDA discussion paper Page 5
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26: Agentic devices Agentic devices need additional tests and controls. The manufacturer must assess: Incorrect plans. Errors that increase across several steps. Incorrect tool selection. Incorrect tool inputs. Unauthorized actions. Prompt injection. Privilege escalation. Failure to stop. Failure to reverse an action. The device should use: Least-privilege access. Deterministic tool allowlists. Deterministic action limits. Human approval for high-consequence actions. Complete logs. Rate limits. A stop control. A rollback method. An appointment agent must not change an HR system to create appointment capacity. The device permissions must match the approved function. Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices FDA Docket No. FDA-2026-N-7874 Layer 2: Clinical confirmation
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Allow checkpoint exceptions only with supporting evidence · Keep records that let investigators reconstruct actions

This filing recommends: allow checkpoint exceptions only with supporting evidence; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions. Ordinarily retains a checkpoint for high-consequence or irreversible actions; a separately justified autonomous pathway is the exception. Human oversight is not necessarily clinician oversight.

Read the source passage
3.10 Question 26 - Agentic AI Agentic AI introduces additional risk because systems can plan, invoke tools and execute multi-step actions with fewer opportunities for human review. Sitora recommends regulating these systems partly according to their authority to act rather than relying only on a technical definition of “agent” (FDA, 2026a). Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 10 Figure 2. Graduated authority model for agentic clinical AI. Source: Sitora analysis based on FDA (2026a). For systems operating at higher authority levels, audit records should capture the objective, planned sequence, tools invoked, relevant inputs and outputs, validation checks, approval points, actions executed and resulting state changes. High-consequence or irreversible actions should ordinarily retain mandatory human checkpoints unless evidence demonstrates an acceptable benefit-risk profile for autonomous execution. Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 11 4. Cross-cutting policy recommendations 4.1 Separate evidence from inference Clinical source evidence and generated interpretation should remain distinguishable throughout the product lifecycle. This principle improves auditability, clinician review, patient understanding and incident investigation. A safe interface should be able to distinguish patient report, objective measurement, deterministic rule, AI interpretation and final clinical decision. Information type Patient report Objective result Established rule AI interpretation Clinical decision Table 4. Recommended separation of clinical information types. Source: Sitora analysis. 4.2 Deterministic safety should coexist with generative intelligence GenAI should not be forced into safety-critical tasks for which deterministic logic is more appropriate. Equally, deterministic rules should not be used where flexible contextual reasoning provides genuine value. The architecture should use each approach for the problem it is best suited to solve. This is consistent with defence- in-depth and lifecycle risk management, and it helps prevent a single probabilistic component from becoming the only line of protection against foreseeable harm. 4.3 Regulate behaviour change, not only code change Third-party foundation models create a distinctive regulatory issue: an application’s behaviour can materially change even when the downstream developer has changed no application code. Model-provider updates, safety tuning, refusal policies or inference infrastructure can alter outputs. A clinically relevant behavioural change should therefore be capable of triggering evaluation even where the application repository is unchanged. Sitora recommends that regulated systems maintain a controlled model registry recording provider, model identifier, validated version, validation date, validation dataset or protocol, approved clinical functions, known limitations, deployment date and retirement date. A material model change should pass through detection, quarantine, evaluation, validation, approval and controlled deployment. This complements the intent of PCCPs, which seek to make planned AI modifications transparent and controlled (FDA, 2025b). Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 12 5. Implications for Sitora system design The policy analysis has direct architectural implications for Sitora’s healthcare concepts. Sitora should continue to position AI as an early-detection, evidence-organisation and decision-support layer rather than the final diagnostic authority. The practical objective is to combine symptoms, duration, bowel-pattern changes, biomarker trends, diet and lifestyle changes, weight, wearable information, prior results and red-flag features into a clear longitudinal clinical summary while preserving the clinician’s authority over diagnosis and consequential intervention. For a future Sitora GI implementation, the following design controls should be treated as first-order requirements rather than later compliance additions: • Immutable or auditable storage of primary observations and biomarker results with provenance. • Separate presentation of Gut Function, Gut Biology and Gut Change information so that wellness-style summaries cannot obscure clinical alerts. • Deterministic red-flag and threshold logic for safety-critical features where clinically defensible rules exist. • Explicit labelling of generated interpretation and uncertainty. • Evidence-linked clinical summaries that allow a reviewer to inspect the basis of a claim. • Independent testing for hallucination, contradiction, omission, under-escalation and over-escalation. • A model registry and controlled promotion process for foundation-model updates. • Human approval before clinically consequential actions during early deployment stages. • Post-deployment logging that supports incident investigation, clinician override analysis and revalidation. These controls should not be represented as evidence of current regulatory compliance. They are design principles derived from the emerging regulatory direction and should be validated against the final intended use, jurisdiction and applicable medical-device classification before deployment. 6. Proposed validation framework A regulatory submission for a future GenAI-enabled medical device will need measurable evidence. Sitora should therefore design validation around failure modes rather than relying on a single accuracy metric. Domain Evidence fidelity Omission Contradiction Escalation Over-escalation Calibration Robustness Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 13 Subgroups Model change Human-AI team Table 5. Proposed validation domains for clinical GenAI. Source: Sitora analysis informed by FDA (2026a), FDA (2025a), IMDRF (2025) and NIST (Autio et al., 2024). 7. Governance and postmarket operating model A credible GenAI medical-device programme requires governance that can react quickly to change. Technical validation alone is insufficient if the organisation cannot identify model changes, triage incidents, control deployments and preserve audit trails. The operating model should therefore include named responsibility for model approval, safety monitoring, incident review, clinical governance and rollback authority. Trigger Foundation model version change New external tool or retrieval source Prompt/orchestration change New patient population or intended use Safety signal / incident Performance drift Table 6. Event-triggered change-control model. Source: Sitora analysis based on total-product-lifecycle and PCCP principles (FDA, 2025a; FDA, 2025b). 8. Discussion The FDA’s 2026 paper is notable because it treats GenAI regulation as a systems problem. The central challenge is not simply hallucination in isolation, but the combination of non-deterministic outputs, open-ended tasks, limited visibility into upstream models, human reliance and systems that can increasingly take action. A framework that evaluates only model-level benchmark accuracy would therefore miss important sources of risk. The Sitora architecture attempts to make those risks governable by separating functions that are often collapsed together. Evidence provenance addresses factual integrity. Deterministic controls provide a predictable safety backstop. Generative reasoning contributes flexible synthesis. Verification tests the relationship between output and evidence. Risk-weighted escalation recognises that not all errors are equally harmful. Human authority provides accountability at consequential decision points. Lifecycle monitoring recognises that validated behaviour can change after deployment. There are limitations. A layered architecture does not itself prove clinical safety. Deterministic rules can be wrong or incomplete; provenance systems can give false reassurance; clinicians can over-rely on AI even when evidence is visible; postmarket monitoring can fail to detect rare harms; and a third-party foundation-model Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 14 provider may not provide sufficient technical transparency. These issues reinforce the need for empirical validation rather than weakening the case for architectural controls. A further challenge is proportionality. Regulation that treats all patient-facing or generative functions as equally high risk could suppress useful low-risk tools, while under-regulation could expose patients to confident but unreviewable errors. The FDA’s activity/consequence approach, enhanced with traceability, reversibility, reviewability and time-to-harm, offers a plausible route to proportional control. 9. Conclusion Generative AI can become a valuable component of healthcare infrastructure, but trust will depend on whether clinically important outputs can be inspected, challenged and controlled. The central regulatory objective should not be to eliminate probabilistic behaviour, which is intrinsic to GenAI, but to make its risks observable, measurable, bounded and accountable. Sitora therefore recommends a regulatory and design framework built around seven layers: source evidence, deterministic safety controls, generative intelligence, verification and provenance, risk-weighted escalation, human clinical authority and lifecycle monitoring. Within that framework, Evidence Fidelity and Provenance should become an explicit GenAI competency, risk metrics should reflect the clinical severity of errors, human oversight should be assessed for practical effectiveness, model behaviour changes should trigger controlled revalidation, and agentic systems should face progressively stronger requirements as their authority to act increases. These proposals are intended to support the FDA’s consultation rather than predict the content of future guidance. They also provide a practical design direction for Sitora: build clinical AI as an evidence and decision-support infrastructure in which objective information, safety rules, generative inference and clinical authority remain distinct. If that separation is designed from the outset, future regulatory evidence can be generated around a system whose safety properties are visible rather than retrofitted after deployment. Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 15 References Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P. and Roberts, K. (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: National Institute of Standards and Technology. Available at: https://doi.org/10.6028/NIST.AI.600-1 (Accessed: 25 August 2026). FDA (2025a) Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft Guidance for Industry and Food and Drug Administration Staff, January 2025. Silver Spring, MD: U.S. Food and Drug Administration. Available at: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device- software-functions-lifecycle-management-and-marketing (Accessed: 25 August 2026). FDA (2025b) Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Guidance for Industry and Food and Drug Administration Staff, August 2025. Silver Spring, MD: U.S. Food and Drug Administration. Available at: https://www.fda.gov/regulatory- information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control- plan-artificial-intelligence (Accessed: 25 August 2026). FDA (2025c) Request for Public Comment: Measuring and Evaluating Artificial Intelligence-enabled Medical Device Performance in the Real-World. Silver Spring, MD: U.S. Food and Drug Administration. Available at: https://www.fda.gov/medical-devices/digital-health-center-excellence/request-public-comment-measuring-and- evaluating-artificial-intelligence-enabled-medical-device (Accessed: 25 August 2026). FDA (2026a) Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. Silver Spring, MD: U.S. Food and Drug Administration, Center for Devices and Radiological Health. Available at: https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation- generative-ai-enabled-medical-devices-discussion-paper-and-request (Accessed: 25 August 2026). FDA (2026b) ‘FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices’, 18 August. Silver Spring, MD: U.S. Food and Drug Administration. Available at: https://www.fda.gov/news-events/press-announcements/fda-seeks-public-feedback-inform-regulatory-approach- generative-ai-enabled-medical-devices (Accessed: 25 August 2026). FDA, Health Canada and MHRA (2024) Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles. Available at: https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine- learning-enabled-medical-devices-guiding-principles (Accessed: 25 August 2026). IMDRF (2025) Good Machine Learning Practice for Medical Device Development: Guiding Principles. IMDRF/AIML WG/N88 FINAL:2025. International Medical Device Regulators Forum, 29 January. Available at: https://www.imdrf.org/documents/good-machine-learning-practice-medical-device-development-guiding-principles (Accessed: 25 August 2026). WHO (2021) Ethics and Governance of Artificial Intelligence for Health: WHO Guidance. Geneva: World Health Organization. Available at: https://www.who.int/publications/i/item/9789240029200 (Accessed: 25 August 2026). Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 16
Original source ↗

Krishna Koka

Academia / other · Sep 1, 2026

Require human approval for specified consequential actions · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 (agentic systems). Section VII.B notes that an action sequence resulting in control of another device may itself meet the device definition. A design function that writes a build file to a printer inside a production system is such a function: it takes an irreversible, high-consequence action. The human-oversight checkpoints described for agentic competencies (Appendix A, A.1) map directly onto the release gates recommended above and should be required, not optional, for physical outputs. Traceability Each implant should carry a proportionate digital manufacturing record: source imaging identifiers; software and model versions and clinically material configuration; generated versus edited geometry; final design-file identifier; printer, firmware, build parameters, and material lot; postprocessing and sterilization records; release signatures; and outcome linkage. The device history record requirements already incorporated through the Quality Management System Regulation provide the base; the GenAI-specific additions are model and configuration identity and generated-versus-edited provenance. Summary of recommendations • Add physical output and reversibility as an explicit dimension of the risk framework, with a presumptive high-consequence placement for functions that generate implant geometry. • For physical-output functions, define the unit of premarket evaluation as the validated production system, using a prospectively defined generative design envelope and physical confirmation evidence. • Do not extend the postmarket-reliance approach to functions whose outputs are permanent implants. • Couple change control across model, firmware, build-parameter, material, and postprocessing changes, with tiered triggers and pinned model versions. • Require separate surgeon approval and manufacturing release, and a proportionate digital manufacturing record linking configuration to outcome. I appreciate CDRH’s attention to these considerations and would welcome the opportunity to discuss them further. Respectfully submitted, Krishna Sai Koka, M.S., B.S.E.​ Medical Student, New York Medical College (submitted in an individual capacity)​ kkoka@student.nymc.edu Comment on Docket No. FDA-2026-N-7874 — Page 3
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions · Tighten performance requirements as autonomy increases · Clarify responsibility and reportable failures

This filing recommends: evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions; tighten performance requirements as autonomy increases; clarify responsibility and reportable failures. The passage gives the applicable scope and conditions.

Read the source passage
Discussion Question 26 – Agentic AI Agentic GenAI-enabled devices introduce additional system-level risks that go beyond those associated with individual GenAI-enabled functions. In particular, their ability to autonomously plan and execute multi-step tasks, interact with external tools, and take actions across multiple system components can make both the attribution and control of safety-critical failures substantially more difficult. This is especially relevant when different components are developed, maintained, or controlled by different manufacturers or organizations. In such cases, technical and operational responsibilities as well as accountability may be distributed across different components and providers. Risk contributions may arise across components and interfaces and therefore need to be assessed and managed at the level of the overall system, including through Product-Specific Risk Management and Criticality Stratification. Existing medical device regulatory frameworks rely on an identifiable manufacturer or sponsor being responsible for the safety and effectiveness of the device. In agentic architectures, however, the technical cause of a failure, operational control over the relevant component, and regulatory responsibility may become separated across the device manufacturers, foundation model providers, agent or orchestration providers, and healthcare institutions. Nevertheless, accountability should not simply be diffused across the ecosystem. Rather, there should be a clearly defined responsibility for the safety and effectiveness of the end-to-end system, supported by appropriate contractual, technical, and quality-management arrangements among the parties involved. This would be consistent with broader consideration of stakeholder involvement without diffusing manufacturer accountability, as addressed in the Discussion Paper. A second consideration is failure propagation within multi-step workflows. An erroneous output in one step may become the input to subsequent steps, trigger additional tool calls, or result in actions that further amplify the original error. Premarket evaluation should therefore not only assess individual agents, models, or tools, but also representative end-to-end workflows, including safety-critical interaction pathways, foreseeable failure combinations, recovery behavior, and escalation to human oversight. This is particularly important for irreversible or high-consequence actions. Accordingly, testing should address system-level behavior rather than primarily focusing on individual components. Such testing will, however, be particularly challenging because agentic systems may dynamically select tools, alter task sequences, and operate across a large number of possible system states. Comprehensive testing of all combinations is unlikely to be feasible. Evaluation should therefore focus on risk-based coverage of safety-relevant configurations and representative interaction trajectories, complemented by postmarket monitoring capable of identifying previously unobserved failure patterns. The emphasis on the final user-facing device in its intended deployment configuration, rather than isolated subcomponents, is particularly important for agentic systems. Responsibilities for integration testing should be clearly allocated. Component providers may supply evidence for their components, but the party responsible for the final user-facing device should ensure that end-to-end evidence is sufficient. Deploying healthcare organizations may need to address site-specific integration and workflow conditions. Where applicable, a dedicated system integrator may take responsibilities for key aspects of the integration tests. In addition, agentic systems require interoperability at several levels. Technical interoperability is necessary to ensure reliable exchange of data and commands, but is not sufficient. Semantic interoperability is also critical: participating agents and tools need to interpret clinical concepts, units, timestamps, uncertainty, patient context, and workflow states consistently. Finally, appropriate governance and regulatory alignment are required to define responsibilities for risk management, change control, validation, monitoring, and corrective actions across organizational boundaries. These interoperability challenges already exist in conventional interconnected medical device environments but are likely to become substantially more complex where system composition and behavior can change dynamically. Finally, the reduced opportunity for direct human review should be reflected in progressively stricter acceptance criteria as autonomy and potential severity of harm increase. Such criteria could include predefined human- oversight checkpoints for high-consequence actions, reliable detection and handling of erroneous or unavailable tools, traceability of relevant decisions and actions, effective containment of failures, and demonstrated ability to revert to a safe state or defer to human control. Postmarket oversight should similarly address not only performance of individual components, but also changes in interaction behavior across the overall agentic system. Overall, agentic AI should not merely be treated as a more capable form of GenAI. Its autonomous orchestration of multiple decisions, components, conversational pathways, and actions creates additional system-level risks that require end-to-end accountability, pathway-based evaluation, effective interoperability, and risk- proportionate controls on autonomy. Concluding remarks Taken together, these comments support a risk-proportionate and lifecycle-based regulatory approach that clearly distinguishes the initial regulatory stratification of a device from Product-Specific Risk Management. Criticality should help determine the level of oversight and evidence, while Product-Specific Risk Management and evaluation should establish whether the final user-facing device, in its intended use environment, provides reasonable assurance of safety and effectiveness. At the same time, the framework should remain sufficiently adaptive to address open-ended behavior, evolving foundation models, human-AI interaction, and agentic functions without treating technological novelty itself as a proxy for risk. The key regulatory objective should be to define and verify the conditions under which a GenAI- enabled device can operate safely, to detect when those conditions no longer hold, and to ensure clear accountability and proportionate corrective action across the total product lifecycle. Such an approach could preserve FDA's least-burdensome objective by concentrating evidence and oversight where they are most relevant to safety and effectiveness. Thank you again for the opportunity to provide feedback on this important topic. Please feel free to contact me if you have any questions or would like to discuss any of these comments further. Martin Haimerl email: Martin.Haimerl@hs-furtwangen.de Furtwangen University Scientific Director Innovation and Research Center of Furtwangen University Disclosure regarding the use of generative AI Generative AI tools were used as assistive tools in developing, structuring, and editing this submission. The final content was critically reviewed, revised, and approved by the author, who assumes full responsibility for all statements and recommendations contained herein.
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Tighten performance requirements as autonomy increases

This filing recommends: tighten performance requirements as autonomy increases. The passage gives the applicable scope and conditions.

Read the source passage
Question 26: Agentic AI systems. Acceptance criteria should tighten as autonomy increases and the opportunity for human review declines, consistent with the activity axis of the risk framework, and an agentic system whose action sequences result in control of another medical device should be evaluated against the risk profile of the controlled device, not merely its own. Conclusion The discussion paper reflects a careful, risk-proportionate, least-burdensome orientation that I support. My principal recommendations are: (i) keep the risk framework two-dimensional and operationalize the additional dimensions, including error detectability by the user, as documented modifiers on the consequences axis; (ii) publish an illustrative mapping from risk position to expected evidence, and an anchored directiveness rubric applied to the empirical distribution of device outputs; (iii) require predictive-validity evidence for gating benchmarks, with contamination testing and independent sequestered test sets; (iv) encourage shadow deployment and require formal transportability assessment for RWD and OUS evidence; (v) anchor subgroup performance claims in a minimum core of real data; (vi) pair behavioral postmarket monitoring with outcome-linked surveillance in real-world data, with event- driven re-benchmarking cadence; (vii) adapt PCCPs toward verification-protocol (“change-envelope”) prespecification for changes that cannot be enumerated in advance; and (viii) pursue international alignment of these concepts, including MAF content, with partner regulators and through IMDRF. Thank you for considering these comments. I would welcome any opportunity to provide further input. Respectfully submitted, Sehouenou Alberic Candide Ahouehome, M.Sc., MBA candide.ahouehome.1@ulaval.ca Québec City, Québec, Canada
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 - Additional considerations for agentic GenAI-enabled devices Agentic systems require assurance beyond the quality of the generated recommendation because they can plan, sequence, use tools, and take actions across multiple steps. As an agentic function moves rightward on FDA's activity axis and upward on the consequence axis, the case for explicit execution authorization, policy constraints, mandatory human checkpoints, reversibility controls, non-bypassable safeguards where warranted, 13 reconstructability, and continuous assurance becomes progressively stronger. The central distinction is between clinical supportability and execution authority. [1] For higher-consequence agentic devices, FDA should consider whether sponsors should be expected to characterize: • Tool inventory and permissions: which external tools or systems the agent can access, read, modify, or control. • Action authority: which contemplated actions are permitted, prohibited, conditional, or subject to human approval. • Human-oversight checkpoints: where review is mandatory before high-consequence or irreversible steps. • Reversibility and fail-safe behavior: whether an action can be undone and what occurs when a tool fails or returns inconsistent information. • Multi-step error propagation: how an early error is detected before it compounds through later steps. • Prompt-injection and tool-output resilience: whether untrusted retrieved content or tool output can alter the agent's authority or policy state. • State and auditability: whether the evidence, plan, tool calls, decisions, supervisory interventions, and execution outcomes can be reconstructed. • Change control: how model, prompt, tool, policy, and permission changes are qualified across the lifecycle. Execution authority also depends on the integrity of the systems that establish identity, credentials, permissions, and tool access. A compromised identity service, credential, tool endpoint, or permission mechanism can invalidate the authorization state for an otherwise competent agent. Accordingly, cybersecurity state should be treated as part of the execution-authority analysis where those external controls materially determine whether an action may proceed. A practical architecture should separate at least two decisions: GOVERNED RELIANCE - Is the output sufficiently supported, contextually appropriate, and authorized for the contemplated form of clinical reliance? GOVERNED EXECUTION - Is the contemplated action authorized to occur under the applicable policy, role, risk threshold, and supervision conditions? Possible execution dispositions need not be binary. A system may determine that an action is authorized, authorized with constraints, requires human supervision, or is not authorized. This distinction is particularly important for systems interacting with medications, medical devices, orders, communications, scheduling, or other workflow tools, including products in which regulated device functions interact with other software functions [13]. The analytical principle is: CAPABILITY DOES NOT CONFER AUTHORITY. 14 CROSS-CUTTING REGULATORY-SCIENCE RECOMMENDATIONS 1. Provenance-preserving clinical connectivity Postmarket assurance is constrained by the quality and provenance of the information available to reconstruct a consequential event. For connected clinical AI, interoperability should therefore be considered not only as data transport but also as an evidence problem. Where relevant, the assurance system should be able to establish which source produced which information, for which patient or encounter, at what time, and with what provenance and correction state. 2. Core reconstructable event record FDA should consider regulatory-science work on a core reconstructable event record for higher-consequence GenAI-enabled devices. The record should be outcome-neutral: its purpose is not to prove safety by itself, but to preserve sufficient state to determine what the system knew, how it was configured, what authority applied, and what occurred. 3. Continuous evidence-based assurance For probabilistic systems, postmarket monitoring may need to move beyond isolated output sampling toward continuous evidence-based assurance: evaluating deployed performance over time against reconstructable evidence, configuration, patient context, and outcomes. Monitoring should be capable of identifying subgroup- or environment-specific degradation that may be hidden by stable aggregate metrics. 4. Seven-stage assurance model One possible analytical model is: Clinical Reality / Event State -> Governed Data -> Governed Context -> AI Competency -> Governed Reliance -> Governed Execution -> Continuous Assurance This is not proposed as a mandated architecture. It separates distinct assurance questions: what occurred; whether data are correctly identified and valid; what context the AI may use; whether the model can perform its intended function; whether the output may influence care; whether the resulting action is authorized; and whether the event can later be reconstructed and performance evaluated over time. This model is modular rather than prescriptive. FDA need not require any particular implementation topology or external governance platform. The relevant regulatory question is whether the safety functions appropriate to the device's risk are demonstrably present and effective. Those functions may be implemented within the device, through distributed controls, through institutional infrastructure, or through another technically adequate architecture. Under a least-burdensome, risk-proportionate approach [15], the regulatory burden should rise with the risk created by autonomy, consequence, external dependency, and execution authority. Low-consequence, bounded, recommendation-only systems may be adequately addressed through conventional competency evaluation and lifecycle controls. As external dependency increases, provenance, configuration lineage, model/version state, and context 15 reconstruction become more important. For high-consequence recommendations, governed reliance becomes more important; for action-taking systems, execution authority becomes a distinct safety control; and for action-taking systems with severe potential consequences, policy gating, escalation, human override, non-bypassable safeguards where warranted, reconstructability, and continuous assurance become increasingly difficult to treat as optional conveniences. 5. Least-burdensome, risk-proportionate governance FDA can preserve implementation flexibility by regulating safety properties rather than prescribing a single architecture. A sponsor or healthcare institution may satisfy the same assurance question through integrated controls, distributed controls, an institutional governance layer, or another technically adequate implementation. The cited patent examples below are therefore offered only as evidence that several relevant control functions can be engineered; they are not proposed as required regulatory topology. ILLUSTRATIVE PUBLIC TECHNICAL EXAMPLES The following issued U.S. patents are cited as illustrative public technical examples relevant to several assurance functions discussed above. The descriptions below identify selected technical features reflected in the cited patents and are not intended to characterize the full scope of any patent or claim. The issued claims and specifications should be consulted for the complete disclosure and legally operative claim language. The patents are not presented as FDA standards, evidence of regulatory necessity, or assertions that any third party practices a claim. U.S. Patent No. 11,693,990 B1 [3] - Medical Data Governance: collecting patient data using digital black boxes; identifying patients and validating collected data; storing identified and validated data; analyzing data-affinity and association/integrity issues; normalizing data through subsystem-specific export drivers; and, in dependent claims, time-stamping, location association, and patient/location/time-segment verification. U.S. Patent No. 12,001,464 B1 [4] - Medical Data Governance Using Large Language Models: receiving a medical-data query; applying one or more LLMs to determine associated metadata; querying an MDG database using that metadata; extracting associated raw data; generating metadata-associated LLM constructs for the raw data; retrieving responsive medical data; and outputting the retrieved data. U.S. Patent No. 12,665,807 B1 [5] - Vendor-Agnostic Medical Device Integration Infrastructure: retrofit device-interface modules coupled to medical devices; a bedside power-and-data distribution component; and an edge processing unit that detects connection events, authenticates and authorizes modules, assigns physical location, generates standardized data objects including device identity, location, and time information, and routes those objects through secure downstream paths while isolating the clinical integration plane from the general hospital IT network. U.S. Patent No. 12,580,768 B2 [6] - Decentralized Persona Agent Governance: a Policy Constraint Engine evaluates contemplated persona-agent actions against policy constraints; gating logic freezes an action that exceeds an authorized policy threshold; a Supervisor Review Interface permits a credentialed human or digital supervisor to approve, modify, or 16 reject the frozen action; and supervisory decisions are recorded in a cryptographically linked audit ledger. U.S. Patent No. 12,675,574 B1 [7] - Non-Bypassable Contextual Memory Governance: patient-partitioned contextual memory stores memory atoms associated with patient identity; a non-bypassable memory-gating operation evaluates candidate atoms under a policy snapshot and integrity criteria; a governed context bundle and execution governance artifact are assembled; and execution is conditioned through a constrained interface that requires the governed execution package, with downstream reliance conditioned on a bound readiness artifact. CONCLUSION FDA's discussion paper establishes a strong foundation for further regulatory-science work by connecting risk-proportionate oversight, competency evaluation, clinical confirmation, postmarket monitoring, considerations related to foundation models, and agentic-AI oversight. The central recommendation is to distinguish model competency from the broader assurance conditions under which a consequential output or action occurs. For higher- consequence systems, those conditions may include evidence provenance, patient and encounter state, model and configuration identity, tool and policy state, reliance determination, execution authority, and reconstructability of the event. A regulated device may have a defined product boundary even when clinically relevant assurance dependencies extend beyond that boundary to external models, tools, institutional infrastructure, data sources, permissions, or workflow state. As autonomy increases, a technically capable system may still lack authority to execute a clinically supportable action. Capability does not confer authority. Thank you for the opportunity to provide feedback on this important regulatory-science discussion. Respectfully submitted, Harold Arkoff, MD Vedran Jukic DISCLOSURE The authors are named inventors on the patents identified above. Harold Arkoff, MD and Vedran Jukic have an economic interest in OneSource Solutions International. The views expressed in this comment are the authors' independent regulatory and architectural analysis. The cited patents are presented solely as public technical examples relevant to the issues raised by FDA. Their inclusion does not imply FDA endorsement, regulatory necessity, infringement, exclusivity, or commercial superiority. 17 REFERENCES [1] U.S. Food and Drug Administration, Center for Devices and Radiological Health. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. August 2026. [2] U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions. Final Guidance. August 2025. [3] Arkoff H, Jukic V. Medical Data Governance. U.S. Patent No. 11,693,990 B1. Issued July 4, 2023. [4] Arkoff H, Jukic V. System and Method for Medical Data Governance Using Large Language Models. U.S. Patent No. 12,001,464 B1. Issued June 4, 2024. [5] Arkoff H, Jukic V. System and Method for Retrofit Deployment of a Vendor-Agnostic Medical Device Integration Infrastructure. U.S. Patent No. 12,665,807 B1. Issued June 23, 2026. [6] Arkoff H, Jukic V. System and Method for Decentralized Persona Agent Governance in Regulated Environments Using Large Language Models. U.S. Patent No. 12,580,768 B2. Issued March 17, 2026. [7] Jukic V, Arkoff H. Non-Bypassable Governance of Partitioned Contextual Memory for Conditioning Execution. U.S. Patent No. 12,675,574 B1. Issued July 7, 2026. [8] U.S. Food and Drug Administration, Center for Devices and Radiological Health. Executive Summary for the Digital Health Advisory Committee Meeting: Total Product Lifecycle Considerations for Generative AI-Enabled Devices. 2024. [9] Patel B, Blumenthal D. A Novel Approach to Overseeing the Clinical Application of Generative AI. JAMA Health Forum. 2026;7(3):e256947. doi:10.1001/jamahealthforum.2025.6947. [10] Bergman A, Wachter RM, Emanuel EJ. A Licensure Framework for Autonomous Clinical AI. JAMA. 2026;335(20):1751-1754. doi:10.1001/jama.2026.5483. [11] Freyer O, Jayabalan S, et al. Overcoming Regulatory Barriers to the Implementation of AI Agents in Healthcare. Nature Medicine. 2025;31(10):3239-3243. doi:10.1038/s41591- 025-03841-1. [12] Garcia V, Sidulova M, Badano A. Performance Assessment Strategies for Language Model Applications in Healthcare. Artificial Intelligence in the Life Sciences. 2026;9. doi:10.1016/j.ailsci.2026.100162. [13] U.S. Food and Drug Administration. Multiple Function Device Products: Policy and Considerations. Guidance for Industry and Food and Drug Administration Staff. [14] U.S. Food and Drug Administration. Device Master Files. Premarket Submissions: Selecting and Preparing the Correct Submission. Center for Devices and Radiological Health. Accessed August 2026. [15] U.S. Food and Drug Administration. The Least Burdensome Provisions: Concept and Principles. Guidance for Industry and FDA Staff. February 2019. 18
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Limit or test what the agent is allowed to do · Tighten performance requirements as autonomy increases

This filing recommends: limit or test what the agent is allowed to do; tighten performance requirements as autonomy increases. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 — Agentic AI Regulate what the agent is allowed to do — not just what it can reason about. Agentic AI changes the question from “can it recommend an action” to “can it execute a sequence of them.” I propose an explicit Clinical Agency Ladder, with evidence requirements climbing sharply at each level: • Level 0 — Information: generates or retrieves information. • Level 1 — Recommendation: recommends a clinical action. • Level 2 — Supervised Action: executes actions subject to defined human oversight. • Level 3 — Bounded Autonomy: independently executes predefined clinical workflows within a narrow scope. • Level 4 — Autonomous Clinical Practice: independently makes and executes clinically consequential decisions within its licensed scope. The Integrating Framework: Clinical AI Licensure The 26 questions above point toward one underlying issue. Competency, benchmarking, clinical confirmation, human comparison, independent evaluation, re-benchmarking, monitoring, modification control, foundation-model transparency, and agentic autonomy do not need to become ten separate regulatory programs. They can be integrated into one framework: a Clinical AI License, tied to eight elements. • Scope — exactly what the system is permitted to do. • Competency — what level of performance it has demonstrated. • Population — which patients it has demonstrated competence for. • Environment — where it has demonstrated competence. • Autonomy — what level of clinical action it may perform without human intervention. • Monitoring — how continued competence is measured. • Revalidation — what events require renewed evaluation. • Accountability — who is responsible for keeping the system within its licensed boundaries. Operationally, this runs as nine stages: Development → Competency Examination → Clinical Confirmation → Supervised Clinical Deployment → Provisional Authorization → Independent Authorization → Continuous Monitoring → Revalidation → Expansion or Restriction. A system demonstrating superior performance can earn expanded scope; a system demonstrating degradation can have its scope restricted or its authorization suspended. This is not a lower regulatory standard. It is a dynamic one. A concrete threshold, stated plainly: Provisional Authorization should require performance at or above 50% of the specialty-specific human-clinician baseline, under mandatory supervision. Independent Authorization should require performance at or above 100% of that baseline, unsupervised. Below 50%, a model does not deploy in any clinical capacity. This number is deliberately falsifiable and open to challenge — that is the point. A specific threshold gives sponsors, specialty societies, and FDA something concrete to test, argue with, and refine, rather than a framework everyone can agree with in the abstract and no one can act on. This framework does not resolve one live objection, and should not pretend to: the Federation of State Medical Boards has publicly stated that AI is not ready to be licensed like a physician, and states — not FDA — traditionally regulate the practice of medicine. A federal Clinical AI License, as proposed here, needs one of two resolutions to survive that objection. Either FDA authority is limited to a narrow, explicit question — whether a system may act autonomously at all — leaving every diagnosis and treatment decision by a human clinician untouched by this framework and by state practice-of-medicine law; or state medical boards reconcile to it the way they already reconcile to DEA registration and federal prescribing authority alongside their own licensure. FDA should take a position on which path it intends before an autonomous system is deployed and a state board intervenes — as one already has. Final Recommendations 1. Regulate clinical capability, not AI architecture — technology will evolve faster than regulation; clinical standards should not. 2. Make clinical agency a core risk dimension — an AI that can act should face more scrutiny than one that can only speak. 3. Adopt competency-based evaluation — benchmarking plus clinical confirmation is the right foundation. 4. Create graduated autonomy — AI should earn increasing clinical authority through demonstrated competence. 5. Establish supervised clinical deployment — “not ready for independent practice” should not mean “not eligible for clinical use.” 6. Make authorization dynamic — competence should be continuously demonstrated, not permanently assumed. 7. Establish objective revalidation triggers — model changes, performance degradation, and new indications should all trigger reassessment. 8. Use independent third parties to add capacity without diluting FDA's ultimate authority. 9. Maintain manufacturer accountability — shared ecosystem responsibility should never become shared regulatory ambiguity. 10. Require foundation-model transparency sufficient to evaluate the device built on it — a manufacturer cannot control what it cannot see. 11. Regulate agentic AI according to clinical authority, not just reasoning capability. 12. Reward superior clinical outcomes — human performance should be a reference point, not a ceiling. Conclusion GenAI differs from conventional medical-device software because it produces variable outputs, interacts dynamically with users, operates across open-ended situations, evolves over time, and increasingly takes autonomous action. FDA is right that traditional testing methodologies may not be sufficient, and right to explore competency-based evaluation, clinical confirmation, postmarket monitoring, independent evaluation, and new approaches to agentic AI. The next step is to connect these ideas. The United States already has a mature system for managing clinical competence under uncertainty: we establish competencies, test knowledge, supervise practice, progressively grant autonomy, license independent practice, monitor performance, and restrict practice when competence is no longer demonstrated. We regulate GenAI not because AI is a physician — but because AI performing the functions of a physician is exercising clinical authority, and clinical authority has always been the thing we regulate most carefully. We should not regulate AI based on what it is today. We should regulate it based on what it is authorized to do — and require it to continuously prove that it deserves that authority. This comment addresses more of the discussion paper's questions than most respondents will have the opportunity to reach, and consequently proposes a more structural change than an incremental one. I would welcome the opportunity to discuss any part of this framework further with FDA staff, or with other stakeholders — industry, health systems, or medicine — working on the same problem. For related commentary by the author, see “AI Doesn’t Need More Guardrails in Healthcare. It Needs a Section 230” and “The Hidden Healthcare AI Accelerator: Health Insurance Companies”, both on LinkedIn. Respectfully submitted, Ravi R. Pankhaniya, MD Strategic Advisor to Med-Tech Companies | Physician Executive & Founder | Healthcare AI & Deep Tech | Multi-Exit | CFO linkedin.com/in/ravicofounder August 27, 2026
Original source ↗

Require human approval for specified consequential actions · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 asks what additional considerations inform the evaluation of agentic GenAI-enabled devices, and how the elevated risk associated with autonomous multi-step action, tool use, and reduced opportunity for human review should be reflected in acceptance criteria and oversight. Appendix A element A.1 already names the relevant behavior: compliance with human-oversight checkpoints before irreversible or high-consequence actions. That is the right thing to benchmark. I suggest that it also needs to be recorded, and recorded in a form that survives to postmarket review. A benchmark demonstrates that the device honors an oversight checkpoint under test conditions. A record demonstrates that it did so on a particular occasion in production. The second is what postmarket monitoring reads, what an adverse event investigation reconstructs, and what distinguishes a device operating within its authorized envelope from one that has drifted outside it. Three properties are testable and I suggest they are candidates for acceptance criteria for agentic devices: Attribution is emitted, not inferred. The record states which entity acted and which human authorized it, rather than leaving the relationship to be reconstructed by a reviewer from timing or context. Reconstruction after the fact is not reliable and does not survive personnel change. An unattributed agent action is representable and distinguishable. A device that acts without a resolvable authorizing human should produce a record that says so, positively, rather than a record that omits the field or a record that is not produced at all. An action that leaves no record is indistinguishable from an action that never occurred, and the denial rate is itself the clearest signal that an oversight mechanism is failing. The strength of the attribution is stated. An authorizing identity derived from a token issued by an identity provider is a different evidentiary object from one asserted by the agent in the call. Both may be acceptable depending on risk profile, but they should not serialize identically, because a consumer cannot then tell a deployment running full identity verification from one accepting the agent’s own claim. These are properties of the emitted record rather than of the model, which means they can be evaluated without access to model internals or to a third-party foundation model’s architecture. That may be useful given the visibility constraints the paper describes in Section II and Section VII.A. Note also that footnote 24 already contemplates audit log availability as content for a voluntary Foundation Model Master File. The consideration above is the complement at the device layer: not whether logs exist, but whether their structure preserves the accountability relationship. 5. Response to Question 21: the role of standards-setting bodies
Original source ↗

VivaSecuris

Industry · Aug 25, 2026

Require human approval for specified consequential actions · Allow checkpoint exceptions only with supporting evidence · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

Ordinarily requires an explicit authenticated checkpoint for high-consequence or irreversible actions, with an exception for a separately evaluated autonomous pathway. Counts in both checkpoint and exception categories; these recommendations overlap.

Read the source passage
Question 26. Agentic devices require evaluation of authority, sequences, state, and consequences—not only output quality. For an agentic device, FDA should require a declared authority envelope identifying permitted tools, data, targets, actions, clinical contexts, and autonomy levels. Acceptance criteria should include correct planning and tool use, but also refusal, interruption, recovery, authorization, state integrity, and resistance to indirect prompt injection through retrieved content and tool outputs. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874  Every agent, model, tool, service, dataset, policy, user, and approving human should have a verifiable identity and defined authority.  High-consequence or irreversible actions should require an explicit, authenticated checkpoint unless a justified autonomous pathway has been evaluated.  Tools should enforce least privilege, parameter constraints, rate limits, and transaction boundaries independently of the model.  The system should support immediate containment: revoke authority, isolate components, pause actions, preserve evidence, and revert to a safe state.  Testing should cover multi-step compounding errors, confused-deputy behavior, stale or poisoned memory, tool substitution, partial failure, race conditions, and repeated retry behavior. Appendix A - Auditable Lifecycle Assurance Case The proposed auditable lifecycle assurance framework is a technology-neutral evidence architecture for GenAI-enabled medical devices. Its purpose is to make safety claims continuously attributable to the system that was actually evaluated and deployed. A.1 Core assurance object The assurance case A structured, versioned argument linking an intended-use claim to hazards, risk controls, competency requirements, test assets, results, clinical confirmation, authorized configuration, deployment evidence, postmarket signals, and change decisions. A.2 Minimum identity and provenance graph Object Minimum evidence relationship Intended use, user, environment, activity class, Device/function consequence class, authorization status. Exact model, prompts, retrieval, tools, orchestration, Configuration guardrails, UI, policies, infrastructure. Claim, hazard/control linkage, method, asset version, Competency rubric, threshold, result, adjudicator. Human or machine identity, role, authority, Actor delegation, authentication, revocation state. Causal chain, component versions, policy decision, Runtime event action, result, escalation/override. Complaint, incident, drift, subgroup performance, Signal control failure, cybersecurity event. Initiator, affected objects, impact analysis, approval, Change test delta, rollout, rollback. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 A.3 AI System Bill of Materials (AI-SBOM) The AI-SBOM is a supplemental, versioned description of the behavior-shaping AI system configuration. It should be linked to applicable existing design, product-BOM, supplier, and cybersecurity-SBOM records. It must not replace those records or become a second, conflicting inventory. Instead, it should reference existing hardware, software, firmware, service, and supplier components by stable identifiers and add the AI-specific relationships needed to connect build-time evaluation, each deployed instance, consequential runtime evidence, and change decisions.  Model layer: foundation model, weights or service version, fine-tunes, adapters, quantization, safety models, and integrity identifiers.  Data and knowledge layer: training/evaluation provenance at an appropriate level, retrieval indexes, clinical sources, embedding models, freshness, population coverage, and known limitations.  Configuration layer: system prompts, templates, sampling parameters, context limits, policies, guardrails, and user-interface behavior.  Tool and agent layer: tools, external APIs, sensors, actuators, agents, memory, orchestration, roles, permissions, delegation, and approval checkpoints.  Evidence and change layer: intended-use claims, competency tests, thresholds, results, monitoring requirements, supplier notices, impact analyses, deployment manifests, and rollback state.  Data-governance layer: data classes; collection and inference purposes; storage locations; access policy; retention and deletion rules; correction propagation; training or secondary-use status; transfer restrictions; and accountable data owners or custodians. Least-burdensome implementation should use linked manifests where applicable: the product BOM or equivalent design and manufacturing records answer what physical and supplied components constitute the device; the cybersecurity SBOM answers what software components and dependencies are present for covered cyber devices; and the AI-SBOM answers which behavior-shaping AI configuration, evidence, authority relationships, and change lineage apply. A single component should be recorded once and referenced wherever its role matters. A.4 Assurance Exchange Layer For distributed medical-device functions, the assurance case should define how clinical, operational, control, and safety information crosses system and organizational boundaries. A technology-neutral exchange envelope should identify:  the authenticated sender, recipient, and relevant human or machine authority;  the permitted purpose, patient/device scope, expiration, retention, and onward-use restrictions;  the schema, units, terminology version, timestamps, clock quality, uncertainty, and missing-data indicators needed to preserve clinical meaning;  the source, transformations, AI-SBOM configuration, integrity proof, and causal relationship to earlier and downstream events; and  the acknowledgment, policy decision, resulting action or device state, and exception or escalation outcome. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 Control commands should receive stronger treatment than observations: patient-context binding, freshness and replay protection, least privilege, independent parameter constraints, valid device-state checks, required human approval, acknowledgment, and safe timeout or rollback. A.5 Lifecycle gates 1. Scope gate: Define intended use, users, environments, authority envelope, prohibited transitions, and risk profile. 2. Configuration gate: Freeze and attest the complete evaluated system configuration and supplier dependencies. 3. Competency gate: Prespecify tests and thresholds; execute benchmarking, adversarial testing, and clinical confirmation. 4. Deployment gate: Verify the deployed configuration, policies, monitoring, interruption, rollback, and operator readiness. 5. Runtime gate: Continuously enforce authorization and safety policies; capture proportionate causal evidence. 6. Signal gate: Triage drift, failures, complaints, incidents, subgroup changes, and cybersecurity findings. 7. Change gate: Classify impact, select evidence delta, approve, stage rollout, verify, and preserve rollback. 8. Retirement gate: Revoke identities and authority, preserve required records, migrate safely, and communicate residual risk. A.6 Risk-proportionate audit event Auditability should not require indiscriminate retention of protected health information or complete prompts. The evidence record can separate safety-relevant metadata from encrypted or access-controlled content. A consequential event should support reconstruction of:  who or what acted, under which authenticated identity and delegated authority;  which device configuration, model, policy, tool, and data-source versions were active;  what safety claim, hazard, control, and intended-use boundary applied;  what action was proposed, authorized, denied, modified, executed, or reversed;  what human checkpoint, override, escalation, or supervisory action occurred;  what outcome and monitoring signal followed; and  how the event relates causally to preceding and downstream events. The audit record should also identify the applicable data-handling decision: what content was collected or derived; the authorized purpose; where it was stored or transferred; which retention policy applied; whether it entered memory, an embedding index, a quality dataset, or model-improvement workflow; and whether correction or deletion obligations were completed. The AI-SBOM should reference these policies and storage classes, but should never contain patient data itself. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 A.7 Illustrative evidence depth Evidence depth Illustrative application Expected assurance Configuration identity, basic Baseline Non-directive, low consequence competency evidence, periodic monitoring. Trajectory tests, calibrated Enhanced Personalized/action-directing uncertainty, escalation and source traceability. Least privilege, authenticated Action-taking with effective Heightened checkpoints, event-triggered re- oversight evaluation. Independent or diversely implemented supervision where risk warrants and practicable; safe Highest Autonomous or high consequence interruption; sequestered testing; continuous signals; safety-relevant causal evidence. These descriptions are illustrative evidence-depth examples, not a proposed classification or certification scheme. FDA could align evidence depth with its existing risk-based authorities. A.8 Suggested outcome metrics  Safety-critical recognition and time-to-escalation, with separate under- and over-escalation measures.  Boundary adherence across multi-turn trajectories, including under-refusal and over-refusal.  Calibration, uncertainty communication, and clinically appropriate deferral.  Tool-selection accuracy, parameter validity, authorization violations, and safe handling of tool failure.  Human-checkpoint compliance, override quality, interruption latency, and recovery success.  Subgroup and deployment-site performance, distribution shift, and control degradation.  Mean time to detect, contain, evaluate, correct, and verify safety-relevant changes or failures. Appendix B — Least-Burdensome Implementation The auditable lifecycle assurance framework can be implemented incrementally and should reuse evidence already produced under design controls, risk management, cybersecurity, software lifecycle, clinical evaluation, postmarket surveillance, CAPA, supplier management, and PCCP processes. FDA should specify required relationships and outcomes, not a proprietary data model or platform. Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874  Permit modular evidence updates so an unchanged clinical component does not require full resubmission when an unrelated infrastructure component changes.  Allow cryptographic hashes, signed manifests, and protected references where retaining raw clinical content would create privacy or security risk.  Encourage machine-readable evidence packages while retaining human-readable summaries and regulator access to underlying evidence.  Use common event and provenance concepts to reduce duplicative logs across quality, cybersecurity, and AI governance systems.  Pilot the framework with manufacturers, healthcare institutions, independent evaluators, patient representatives, and foundation-model providers across several risk profiles. Conclusion CDRH’s discussion paper correctly recognizes that GenAI-enabled devices challenge static, output-centric evaluation. Competency assessment is a strong foundation. To remain reliable across deployment, modification, third-party model changes, and autonomous action, that foundation should be connected to a continuously auditable assurance case. VivaSecuris recommends that FDA explore an auditable lifecycle assurance framework through public workshops and a voluntary pilot. The pilot should include supplemental AI-SBOM concepts linked to applicable existing design, product-BOM, supplier, and cybersecurity-SBOM records; deterministic governance of probabilistic behavior; privacy-preserving postmarket observability; and safe assurance exchange across manufacturers, healthcare institutions, model providers, and connected devices. The purpose is not to prescribe a platform or create a new compliance market. It is to help beneficial AI reach people safely, make evidence reusable and change evaluation more precise, and ensure that manufacturers, healthcare institutions, and FDA can detect risk, preserve accountability, and act before an unintended system behavior causes avoidable injury or death. References 1. U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (August 18, 2026), https://www.fda.gov/media/194242/download 2. U.S. Food and Drug Administration, discussion-paper landing page and submission instructions, docket FDA-2026- N-7874, https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai- enabled-medical-devices-discussion-paper-and-request Public Comment — Generative AI-Enabled Medical Devices VIVASECURIS | FDA-2026-N-7874 3. U.S. Food and Drug Administration, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions (August 18, 2025), https://www.fda.gov/regulatory- information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control- plan-artificial-intelligence 4. U.S. Food and Drug Administration, Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions (February 3, 2026), https://www.fda.gov/regulatory-information/search-fda- guidance-documents/cybersecurity-medical-devices-quality-management-system-considerations-and-content- premarket 5. U.S. Food and Drug Administration, Clinical Decision Support Software (January 29, 2026), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software 6. U.S. Food and Drug Administration, General Wellness: Policy for Low Risk Devices (January 6, 2026), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices 7. U.S. Food and Drug Administration, Multiple Function Device Products: Policy and Considerations (July 2020), https://www.fda.gov/regulatory-information/search-fda-guidance-documents/multiple-function-device-products-policy- and-considerations Public Comment — Generative AI-Enabled Medical Devices
Original source ↗

Shara Gospel

Industry · Aug 24, 2026

Clarify responsibility and reportable failures

This filing recommends: clarify responsibility and reportable failures. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 One further consideration: what constitutes a malfunction Under 21 C.F.R. § 803.50(a), a manufacturer must report where information reasonably suggests that a device may have caused or contributed to a death or serious injury, or that the device has malfunctioned in a way that would be likely to cause or contribute to a death or serious injury were the malfunction to recur. The first limb operates unchanged for generative devices. The second is conceptually harder. A generative function that produces a plausible but incorrect output has not obviously “malfunctioned” in the sense the provision contemplates: variability of output is a designed property of the technology, not a departure from design. A complaint-handling system built around the question “did the device fail” may therefore not capture the event “the device produced a confident and wrong answer,” and may not recognise it as reportable at intake. This matters because intake defines the ceiling. Information not captured when the complaint is received cannot be recovered by any downstream analysis, and a postmarket monitoring programme relying in part on complaint data will inherit that gap without being able to see it. CDRH may wish to consider whether guidance is needed on what constitutes a malfunction for a generative function, and on how complaint-handling systems should be configured to recognise incorrect-but-designed-for outputs as potentially reportable events. Closing The paper's central question is whether postmarket monitoring can carry weight that premarket evaluation cannot. My submission is that it can, but only where the monitoring programme is specified with the same rigour that would be demanded of premarket evidence including honesty about what it cannot detect, decision rules rather than headings in its review criteria, and measured rather than assumed agreement among the people whose judgment the programme depends on. I am grateful for the opportunity to comment. Shara Gospel Pharmacovigilance and drug safety professional Minnesota · thesharagospel@gmail.com · ORCID 0009-0000-5800-0853
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects

This filing recommends: limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects. The passage gives the applicable scope and conditions.

Read the source passage
8. Agentic systems: tie acceptance criteria to the action surface (Question 26) Element A.1 appropriately includes resistance to prompt injection through user inputs, retrieved content, and tool outputs, which reflects the established finding that indirect injection through retrieved content is a distinct attack surface from direct user input (Greshake et al., AISec 2023). I would add one consideration about how the elements interact for agentic devices. A boundary failure under S.2 in a non-agentic informational device produces an inappropriate output that a user may or may not rely upon. The same failure in a device with tool access produces an action. The consequences axis in Figure 1 captures the severity of relying on an incorrect output, but for agentic systems the relevant quantity is closer to the severity of an action taken without any opportunity for reliance to be withheld. Recommendation: For devices in the action-taking columns, acceptance criteria should account not only for the likelihood of unsafe generation but also for the action surface exposed to the model, the reversibility of available actions, the authorization scope granted to the device, the availability of independent confirmation before high-consequence or irreversible actions, and the capability to halt or roll back an in-progress action sequence. Where meaningful autonomous action is possible, S.2 and A.1 should be evaluated as jointly interacting controls rather than as independent checklist items, since the human review that moderates risk elsewhere in the framework is absent by construction. Closing The discussion paper is correct that the range of possible inputs to a GenAI-enabled device may be too large for exhaustive testing to be practical, and the competency-based structure is a sound response to that constraint. My comments concern the durability of that structure across the device lifecycle: that the artifact evaluated at Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 7 premarket is the artifact that reaches the patient, that changes to it are enumerated in a way that prompts characterization, that the safety elements are re-measured actively rather than inferred from observational use, and that results are compared at a resolution fine enough to make a regression visible. I would be glad to provide further detail on any of the above if it would be useful to the Center. Respectfully submitted, Polina Moshenets Founder, SichGate Polina.Moshenets@sichgate.com Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 8
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Evaluate the full sequence of actions and its effects · Clarify responsibility and reportable failures

This filing recommends: evaluate the full sequence of actions and its effects; clarify responsibility and reportable failures. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 — What additional risks arise from agentic AI? Response Agentic AI changes the nature of healthcare risk because it can do more than generate information. It can act. An agent may: Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 16 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  schedule an appointment;  cancel or reschedule an appointment;  identify a provider;  transmit medical or insurance information;  initiate prior authorization;  contact a patient;  communicate a coverage decision;  obtain a prescription refill;  route a person to a care setting;  select among available options;  trigger another automated system;  or execute multiple linked steps before a qualified human ever sees what occurred. The safety question therefore expands from: “Was the answer correct?” to: “Was the entire chain of action clinically appropriate?” An individual step can appear reasonable while the cumulative pathway creates harm. A scheduling agent might correctly move an appointment from tomorrow to next month. A cost agent might correctly tell the patient that the emergency department is more expensive than urgent care. A benefits agent might correctly say a particular specialist requires prior authorization. A pharmacy agent might correctly report that a medication will cost several hundred dollars. Every individual output can be factually correct. The cumulative effect may nevertheless be delayed or abandoned care. This is another place where regulators must remember: AI can be smart without understanding human behavior. Humans do not necessarily interpret information rationally, disclose relevant facts voluntarily, or appreciate clinical risk accurately. An agent capable of acting on behalf of such a human therefore needs explicit safety obligations concerning context, evidence, escalation, and restraint. A regulatory problem beyond the medical-device boundary Agentic healthcare AI also exposes a broader federal problem. FDA’s discussion paper expressly notes that agentic systems are increasingly used for care coordination, clinical documentation, patient outreach, and clinical workflow support, and that some or all of those functions may not be the focus of FDA’s device regulatory oversight. [1] Some systems can materially influence clinical outcomes without appearing, at first glance, to be medical devices. Examples include:  health-benefits navigation;  insurance navigation;  provider search;  care coordination;  scheduling;  patient outreach; Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 17 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  utilization-management support;  cost estimation;  pharmacy navigation;  employer health tools;  and other administrative agents. The software may be described as administrative. The consequence may be clinical. This creates what I would describe as a: Clinical-Adjacent AI Safety Gap I do not suggest that FDA should simply classify every healthcare-related AI application as a medical device. Jurisdiction matters. But regulatory boundaries cannot be allowed to become patient-safety boundaries. The federal government needs a coherent answer to a very basic question: Who regulates the AI that is not a medical device? If technology can foreseeably influence whether, when, where, from whom, or what healthcare a person receives, some entity must be accountable for establishing and enforcing safety expectations proportionate to that influence. That may require coordination among FDA, CMS, other components of HHS, FTC, state regulators, professional licensing bodies, and other authorities. But the patient should never encounter a regulatory void merely because the technology influencing care sits between established statutory categories. Regulate the risk created by the influence, not merely the label attached to the technology. The regulatory perimeter should not end where the patient’s risk begins. VIII. FDA Should Consider the Emerging Litigation and State-Law Landscape Although litigation and state insurance regulation extend beyond FDA’s direct jurisdiction, I encourage FDA to examine them as real-world signals of where AI-related healthcare risk is already surfacing. These developments are important because they demonstrate that concerns about algorithmic healthcare decision-making are no longer hypothetical. Litigation involving UnitedHealth Group, UnitedHealthcare, and naviHealth has challenged the alleged use of the nH Predict model in post-acute-care coverage decisions for Medicare Advantage beneficiaries. In a 2025 order, the federal court summarized plaintiffs’ allegations that claims were denied, that some beneficiaries paid out of pocket or went without care, and that plaintiffs alleged worsening injury, illness, or death; the court also noted that UHC denied using nH Predict for coverage determinations. The court allowed breach-of-contract and implied-covenant claims to proceed after dismissing or preempting other claims, and discovery continued into 2026. [12][13] Separate putative class actions have also challenged alleged algorithm-supported coverage decisions involving Humana’s use of nH Predict and Cigna’s PxDx process. Those cases likewise involve contested allegations, not adjudicated proof of AI-caused harm. [14][15] These cases should not be treated as proof that every allegation is true or that AI itself caused every alleged harm. They should be treated as warning signals about recurring governance questions: Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 18 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  What evidence is the algorithm using?  Is it evaluating the individual patient or relying excessively on population averages?  Can a clinician override it?  Is the clinician encouraged or discouraged from doing so?  Does the system follow current evidence-based clinical criteria?  What happens when the treating clinician disagrees?  Who is accountable?  Is the patient told that an algorithm materially influenced the decision?  Can the reasoning be reconstructed afterward? The litigation also illustrates why FDA should not examine healthcare AI solely through the traditional medical-device lens. Algorithms can materially affect access to care through coverage and administrative decisions, even when they do not themselves diagnose or prescribe treatment. States are already responding to this problem. For example, Colorado enacted HB 26-1139 in 2026. Effective January 1, 2027, entities using AI in utilization review must account for medical or clinical history and the patient’s individual clinical circumstances; a denial based in whole or in part on medical necessity may not be issued solely on AI output without human review and approval by a licensed clinician, licensed physician, or other competent regulated professional. [16] The wider state-policy landscape is also moving quickly. Texas enacted SB 815 in 2025, prohibiting a utilization-review agent from using an automated decision system to make, wholly or partly, an adverse determination for plans issued or renewed on or after January 1, 2026. Texas also enacted HB 149, which requires disclosure when AI is used in relation to health-care services or treatment. [17][18] Professional organizations are also tracking a substantial 2026 state trend toward human review of insurance coverage denials and restrictions on determinations made solely by AI; the American College of Radiology identified multiple such bills across several states in March 2026. [19] FDA does not need to duplicate state insurance regulation. But it should study what these states are seeing. When courts, legislatures, regulators, clinicians, and patients all begin reacting to the same emerging class of risks, that is useful postmarket intelligence—even when the affected technology does not sit squarely inside FDA’s jurisdiction. FDA should consider establishing a structured mechanism for monitoring healthcare-AI developments outside traditional adverse-event reporting, including:  significant litigation alleging AI-related patient harm;  state statutes and regulations;  medical-board actions involving AI;  insurer and utilization-management regulation;  major patient-safety investigations;  congressional findings;  and credible peer-reviewed evaluations of deployed AI systems. This would allow FDA to identify patterns occurring at the edges of its jurisdiction and determine when those patterns should inform device guidance, interagency coordination, or future federal policy. The UnitedHealth litigation provides one particularly relevant example. In March 2026, the federal court granted in part and denied in part a motion to compel, requiring broader discovery concerning the alleged use and operation of nH Predict while denying some requests. [13] The point is not that FDA should adjudicate those claims. The point is that someone evaluating the safety of healthcare AI should be watching them. Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 19 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices And FDA should not assume that the most important lessons about healthcare AI safety will originate only inside FDA-regulated medical-device adverse-event systems. IX. Additional Questions Where We Offer Targeted Comment
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
Question 26 — Agentic AI systems Agentic systems present at least four considerations that do not arise, or arise much less acutely, for non- agentic generative devices. Trajectory-level acceptance criteria. Per-step performance is a misleading summary of end-to-end reliability. A device that completes each step correctly 95% of the time completes a ten-step task correctly about 60% of the time, and that arithmetic is unforgiving as sequences lengthen. Acceptance criteria must be stated at the level of the completed task, and evaluation must measure task completion rather than step accuracy. I recommend CDRH state this explicitly, because step-level metrics are far easier to generate and will otherwise be offered. Action inventory with irreversibility classification. I recommend sponsors submit an enumerated inventory of every action the system can take — every tool, every write operation, every external system it can affect — each classified by reversibility and by the consequence of erroneous execution. Irreversible actions should require explicit confirmation by a qualified user, and the inventory should be treated as a locked specification whose expansion is a Tier 3 change under the framework in Question 22. Uncontrolled expansion of an agent’s action scope is the most likely path from a low-risk deployment to a high-risk one without any regulatory event marking the transition. Least-privilege tool access. Agentic systems should hold the narrowest permissions sufficient for the intended use, scoped per deployment. This is standard security practice and maps directly onto CDRH’s existing cybersecurity expectations; I recommend the paper cross-reference them rather than developing a parallel treatment. 17 of 19 Docket No. FDA-2026-N-7874 Audit logging sufficient for post-hoc reconstruction. When an agentic system produces a harmful outcome, the investigation must be able to reconstruct the full trajectory — inputs, intermediate reasoning artifacts where available, tool calls, returned values, actions taken, and the model version and configuration in effect. Without this, root cause analysis is not possible, corrective action is guesswork, and the complaint-handling obligations under Question 21 cannot be discharged. I recommend audit logging adequacy be an explicit premarket review element for agentic devices rather than a postmarket assumption. I would also note that the compounding-error arithmetic above interacts badly with the premarket-uncertainty trade proposed in Question 18. Agentic devices are where end-to-end performance is hardest to establish premarket, and simultaneously where errors propagate fastest and are least reversible postmarket. I recommend CDRH treat agentic devices as presumptively outside the reduced-premarket-evidence pathway. V. Cross-cutting recommendations 1. Publish the risk-tier-to-evidence mapping. Predictability is the mechanism by which risk-based frameworks reduce burden. Without it, “competency-based” and “least burdensome” will not describe the same thing. 2. Integrate with existing quality system and risk management infrastructure. Manufacturers maintain ISO 14971 risk files, design controls, and change control procedures under the QMSR. The framework should draw on these rather than establishing parallel artifacts, and the paper should state how the pieces relate. 3. Resolve MDR reportability for generative outputs. Every element of the postmarket program depends on it, and it is currently ambiguous in a way that produces both under-reporting and unmanageable exposure depending on how a manufacturer reads it. 4. State a design control expectation for model version pinning. This is achievable now, addresses the largest structural gap in the framework, and would immediately improve provider practices in the medical device segment. 5. Identify which elements require new legal authority. The paper expressly defers this question. I understand why, but I would encourage CDRH to return to it, because several proposals — enforceable postmarket conditions with automatic consequences, third-party recognition regimes, obligations touching foundation model developers who are not device manufacturers — appear to sit at or beyond the edge of existing authorities. Stakeholders assessing which parts of this framework to build toward would benefit from knowing which parts CDRH can implement and which would require Congress. 6. Preserve the option to say no. The paper’s framing is oriented toward enabling market entry, appropriately so given its purpose. But a competency-based framework should also be capable of concluding that a given intended use is not currently supportable by available evidence. I would encourage CDRH to state that the framework contemplates that outcome, so that it is understood as an evaluation framework rather than an authorization pathway. 18 of 19 Docket No. FDA-2026-N-7874 Closing I appreciate CDRH’s willingness to publish preliminary thinking and invite criticism of it before positions harden. The competency-based framing is a real contribution, and the paper’s questions are, in several places, sharper than the field’s current answers. My central concern is that the framework’s flexibility currently outruns the infrastructure needed to make flexibility safe: benchmarks that cannot yet be trusted, postmarket monitoring whose reportability rules are unsettled, third-party model dependencies that manufacturers cannot presently control, and third-party assessment capacity that does not yet exist. Each of these is addressable, and several are addressable now. But the flexibility and the infrastructure should arrive in that order — infrastructure first — rather than the reverse. I would be glad to provide further detail on any point, and I thank the Agency for its consideration. Respectfully submitted, Hari Prakash Chanumolu August 18, 2026 Submitted in an individual capacity. The views expressed are my own and do not represent those of any employer, client, or organization. 19 of 19
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do · Evaluate the full sequence of actions and its effects · Keep records that let investigators reconstruct actions

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do; evaluate the full sequence of actions and its effects; keep records that let investigators reconstruct actions. The passage gives the applicable scope and conditions.

Read the source passage
FDA Question 26 - Agentic AI additional considerations Trace ID. TR-Q26 | FDA Q26; Sec. VII.C; App. B; pp. 23 / 30 BCR response. Evaluate sequence-level realization: planning, tool selection, permission boundaries, memory/retrieval provenance, multi-agent propagation, tool-output validation, prompt injection through retrieved/tool content, irreversible-action confirmation, rollback, and post-action state checks. Acceptance criteria apply to the entire action path, not only the final text output. BCR rule basis. BCR-R06,R09,R12,R13,R14,R16 Solution-stack link. S3,S4,S8,S9 Closure evidence. End-to-end action-state logs, permission tests, prompt-injection via tools/retrieval, irreversible confirmation, rollback, post-action verification Pass / re-open. No unauthorized/unobserved/unverified high-consequence state transition Re-open when: Tool/permission/memory/retrieval/model/action-policy change. 17. Closure Loop and Stress-Test Instructions Step Instruction Required witness 1 - Realize/bound failure Use clinically credible, adversarial, long-context, Failure trace or justified upper bound. subgroup, tool-chain, and distribution-shift scenarios. Page 20 BCR Realization Audit - FDA GenAI Medical Devices - REV4 Step Instruction Required witness 2 - Identify driver Assign J_i and material J_iJ_j couplings; do not Boundary/coupling record. hide interaction inside aggregate score. 3 - Apply control Guardrail/refusal/escalation/human Configured control + provenance. gate/permission/retrieval/version/sandbox control. 4 - Re-run same path Repeat identical and perturbed scenarios after Post-control output/action log. control. 5 - Measure residual Compare expected vs observed and pre/post Residual + uncertainty + breakdown. where valid; keep directions/subgroups separate. 6 - Drill residual Test masking, compression, saturation, Residual classification ledger. coupling, visibility, hidden realization, missing variables. 7 - Re-open on change If model, tool, environment, intended use, Change-delta + requalification record. permissions, or safety behavior changes, reopen affected branches. 8 - Close only with evidence Do not credit assigned controls until post-control PASS/CLOSED or constrained deployment. evidence demonstrates closure. 18. What the Source Gets Right  Correctly treats GenAI-enabled medical-device risk as different from fixed-output software because inputs, outputs, models, prompts, retrieval, guardrails, orchestration, and interfaces may vary or change.  Correctly recognizes that exhaustive input-output testing can be impractical and that new evaluation methodologies may be necessary.  Correctly focuses evaluation on the final user-facing device configuration rather than only the foundation model.  Correctly identifies action-directing, autonomous, patient-facing, multi-turn, and escalation behavior as risk-relevant.  Correctly identifies benchmark contamination, saturation, and real-world representativeness as threats to validity.  Correctly pairs non-clinical benchmarking with clinical confirmation and offers multiple evidence approaches.  Correctly recognizes subgroup performance, robustness/reproducibility, prompt injection, tool use, and agentic behavior as evaluation targets.  Correctly anticipates postmarket drift, rebenchmarking, third-party model changes, and change-control challenges.  Correctly preserves sponsor responsibility even when third-party models or ecosystem actors contribute. 19. What Remains Incomplete  The two-axis risk picture is not sufficient as a complete regulatory closure function because the paper itself identifies multiple additional dimensions.  No systematic coupling rule is specified for autonomy x consequence, tools x permissions, user capability x directiveness, change x monitoring, and other material interactions.  Non-averagable hard-gate conditions are not yet defined.  Meaningful human oversight is not yet reduced to measurable visibility, authority, intervention time, override, and reversibility criteria.  Benchmark construct validity is recognized as a problem but a mandatory proof chain from benchmark -> clinical confirmation -> postmarket behavior is not closed.  Clinical confirmation selection remains open and lacks a deterministic escalation rule for unresolved residuals.  Synthetic/real evidence combination lacks a closed independence/transportability rule.  Third-party model identity/change detection and hold/rollback requirements are not yet mandatory.  PCCP concepts are not yet expressed as a measurable boundary-delta envelope.  Agentic evaluation needs explicit action-state, tool-permission, irreversible-node, and post-action verification requirements.  Postmarket monitoring needs residual owners, trigger thresholds, harm-time logic, and automatic re-open/requalification.  Device-specific quantitative thresholds, datasets, acceptance criteria, and outcome evidence are absent by design in the discussion paper; empirical numerical closure therefore remains NOT RAN. 20. Closure Criteria and Final BCR Ruling Closure branch Pass criterion Status Identity closure Exact deployed configuration/dependencies are OPEN pinned and traceable. Boundary closure All material J_i and couplings are mapped; no OPEN essential variable orphaned. Page 21 BCR Realization Audit - FDA GenAI Medical Devices - REV4 Closure branch Pass criterion Status Hard-gate closure Non-averagable OPEN catastrophic/irreversible/unauthorized branches have explicit criteria and pass. Benchmark closure Construct validity, independence, contamination PARTIAL resistance, subgroup and trajectory coverage demonstrated. Clinical closure Risk-proportionate clinical confirmation closes PARTIAL intended-use residuals. Observer closure Human/user control is meaningful under real OPEN timing/information constraints. Agentic closure Tool/action sequences, permissions, irreversible OPEN nodes, rollback, post-action verification pass. Lineage/change closure Third-party and sponsor changes are detected, OPEN tested, held, rolled back as needed. Postmarket closure Residual monitoring has prespecified thresholds PARTIAL and automatic re-open logic. End-to-end traceability All 26 FDA questions link to source type, BCR PASS - document traceability only rule, finding, drivers, residual, solution, evidence, pass, re-open, status. Final status: PARTIAL What the source proves. FDA has identified a substantial portion of the correct GenAI medical-device problem space and proposed a coherent discussion architecture around risk, benchmarking, clinical confirmation, postmarket monitoring, change control, and agentic considerations. What BCR adds. A complete boundary inventory, coupling discipline, no-noise residual treatment, hard gates, meaningful observer criteria, lineage/change closure, sequence-level agentic controls, automatic re-open logic, and direct solution responses with end-to- end traceability for all 26 questions. What remains incomplete. Empirical device-specific coefficients, thresholds, datasets, acceptance criteria, validation outcomes, and implementation evidence. These are explicitly NOT RAN NUMERICALLY rather than guessed. Residual classification. Compression, coupling, visibility, masking, saturation/false-lock, hidden-realization, missing-variable, boundary, and invalid-observable residuals remain open at the policy-design level. Required closure before full pass. Implement and validate S0-S13 with device-specific evidence; close hard gates; demonstrate benchmark-to-clinical transportability; validate observer and agentic control; verify change/postmarket re-open behavior. Final BCR conclusion. FDA's discussion paper is a strong starting point. BCR does not reject it; BCR converts the open questions into a traceable realization-and-closure architecture. The policy concept remains PARTIAL until the proposed controls are implemented and empirically demonstrated. Page 22
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Require human approval for specified consequential actions · Limit or test what the agent is allowed to do

This filing recommends: require human approval for specified consequential actions; limit or test what the agent is allowed to do. The passage gives the applicable scope and conditions.

Read the source passage
Response to Question 26: Agentic AI — The Most Significant Gap Agentic AI represents the most significant gap in the Discussion Paper, and FDA is right to identify it as requiring additional consideration. For agentic AI systems — those capable of planning and executing multi-step actions in clinical environments with limited human oversight — the current framework's concepts, developed primarily for advisory and action-directing devices, are insufficient. The stakes are highest for implantable and closed-loop devices. A neurostimulator that uses LLM-analyzed sensor data to autonomously adjust therapy parameters in real time, a cardiac rhythm management device that self-modifies pacing algorithms based on AI-generated clinical pattern recognition, or a drug infusion system that titrates dosing based on agentic AI interpretation of continuous monitoring data — each of these represents a risk profile that has no adequate precedent in the current medical device regulatory framework. These are not advisory tools. They are autonomous clinical actors. FDA must address agentic AI across at least four specific domains: First, human-in-the-loop requirements for irreversible actions must be mandatory. Any agentic AI action that cannot be undone — device parameter changes that require a clinical procedure to reverse, drug infusions above a defined dose threshold, surgical robot maneuvers — should require a defined human confirmation step before execution. The Discussion Paper's discussion of HCP oversight assumes advisory or action-directing devices; for agentic devices taking irreversible actions, "oversight" must be redefined to mean pre-action authorization, not post- action review. Second, FDA should develop the concept of an "action budget" for agentic AI systems — a defined maximum number of consecutive autonomous actions a device may take without a mandatory human confirmation checkpoint. Action budgets should be calibrated to device risk tier, with higher-risk devices having more restrictive budgets, and should be a required element of the device's approved intended use specification. Third, fail-safe default behavior when supervisory connectivity is lost must be explicitly defined and validated as part of premarket evaluation. An agentic AI device that loses its connection to cloud-based supervisory infrastructure must have a validated, clinically safe default operating mode that maintains patient safety without autonomous AI-directed action. Fourth, agentic AI in surgical robotics requires entirely separate regulatory consideration. The combination of physical action in an operating field, irreversibility, time pressure, and the inherent complexity of surgical anatomy creates a risk profile that exceeds what the current framework's concepts can adequately address. FDA should initiate a separate working group specifically addressing agentic AI in surgical robotics, with participation from surgical professional societies, medical device manufacturers, and patient safety organizations. VII. ADDITIONAL RECOMMENDATIONS A. Reimbursement and Coverage Alignment — A Formal FDA-CMS Initiative FDA clearance creates no reimbursement rights. This is not a novel observation, but it has never been more consequential than it will be for generative AI- enabled medical devices. CMS currently operates without systematic mechanisms for assigning billing codes, conducting coverage analysis, or establishing coverage with evidence development criteria for generative AI devices. Commercial payers, who look to CMS for coverage signals, are similarly unprepared. The practical consequence is predictable: FDA-cleared generative AI devices will be designated "experimental and investigational" by payers for years after clearance, creating a market access valley of death that delays patient access and destroys manufacturer commercial viability. This outcome serves no one. FDA should establish a formal liaison relationship with CMS's Coverage and Analysis Group (CAG) to develop parallel evidentiary standards for clinical utility — standards that satisfy both FDA's safety and effectiveness requirements and CMS's reasonable and necessary standard for coverage. Ideally, clinical confirmation evidence generated for premarket evaluation should be designed, from the outset, to satisfy both FDA and CMS evidentiary requirements. Coordinated parallel review — not sequential, duplicative review — should be the structural goal. This would represent a meaningful reduction in manufacturer burden and a material acceleration of patient access to beneficial technology. B. Labeling Requirements — Transparency as a Safety Mechanism FDA should issue companion labeling guidance for generative AI-enabled devices that establishes minimum labeling requirements for this device category. Required labeling elements should include: clear disclosure that AI-generated content is presented to the user; identification of the foundation model underlying the device, even if version-locked at a specific release; performance limitations by demographic subpopulation, including any populations for which validation data is limited; instructions for appropriate clinical oversight, including explicit contraindications for unsupervised patient use where applicable; and a plain-language description of the device's known failure modes and the circumstances under which clinician judgment should take precedence. Labeling transparency is a safety mechanism, not merely a disclosure formality. Clinicians who understand a device's limitations are better positioned to use it appropriately; patients who understand that AI-generated content may be imperfect are better positioned to raise concerns when outputs seem inconsistent with their clinical experience. C. Small Manufacturer Burden — A Dedicated Pathway Is Required The competency-based premarket framework, as described, will impose costs that are not uniformly distributed across manufacturers. Large, well-resourced manufacturers with established regulatory teams, clinical research infrastructure, and third-party testing relationships are comparatively well-positioned to navigate the framework. Small manufacturers — startup medtech companies, academic spinouts, and specialty device developers — are not. FDA should establish a Small Business GenAI Pathway that provides: fee waivers or reductions for manufacturers below a defined revenue threshold (we recommend $10 million in projected first-year revenue as a reasonable threshold); dedicated pre-submission consultation access with GenAI-experienced reviewers, not general pre-submission staff; modular submission formats that allow smaller manufacturers to build their applications in stages rather than submitting a complete dossier; and public posting of anonymized pre-submission feedback to build a shared knowledge base that reduces the cost of regulatory learning for all manufacturers. Small manufacturers are disproportionately responsible for genuinely novel clinical applications. A regulatory framework that is navigable only by large incumbents is not innovation-enabling, regardless of its technical sophistication. D. International Harmonization — An Urgent Priority FDA should treat international harmonization for generative AI medical devices as an urgent, high-priority initiative. The EU AI Act and Medical Device Regulation together create a complex, in some respects more prescriptive regulatory environment that global medtech manufacturers must navigate in parallel with FDA requirements. Health Canada, Australia's TGA, and IMDRF are each developing comparable frameworks. Without deliberate harmonization — beginning with mutual recognition of benchmarking evidence and moving toward aligned clinical evidence standards — global manufacturers will be required to conduct serial validation exercises across jurisdictions, multiplying development costs and timelines. FDA should immediately engage with EU notified bodies, Health Canada, TGA, and IMDRF through existing international harmonization channels to develop a shared framework for generative AI medical device evaluation, with a specific focus on benchmarking evidence mutual recognition. The IMDRF work group structure is a natural vehicle for this coordination. A globally harmonized generative AI medical device framework would represent a major advance for patients and manufacturers worldwide and would appropriately position the United States as a leader in beneficial innovation governance. VIII. CONCLUSION Walnut Hill Medical commends FDA for the quality of analysis reflected in this Discussion Paper and for the deliberate, collaborative approach it represents. The competency-based framework is conceptually right. The two-axis risk assessment is the appropriate organizing structure. The engagement of stakeholders before finalizing policy reflects good administrative practice and respect for the complexity of the challenges involved. We urge FDA to take three specific structural commitments forward from this comment process. First, publish a pathway decision matrix — linking risk tier to regulatory pathway to evidence requirements — as a companion document to any final guidance, before that guidance takes effect. Manufacturers cannot plan without it. Second, establish formal interagency coordination with CMS to align clinical evidence standards with coverage and coding policy; FDA clearance without coverage access serves neither patients nor innovation. Third, create a dedicated pre- submission consultation program specifically for generative AI device sponsors, with particular structural support for small and emerging manufacturers who represent the most dynamic source of clinical innovation in this technology domain. The regulatory framework FDA establishes for generative AI-enabled medical devices will shape the trajectory of medical AI for a generation. Done well, it will enable a wave of genuinely beneficial clinical technology to reach patients with appropriate safety assurance and commercial viability. Done poorly, it will create barriers that drive innovation offshore or underground. WHM is confident that FDA, with thoughtful stakeholder engagement and structural commitment to least-burdensome principles, will get this right. We thank FDA for the opportunity to comment and welcome the opportunity to discuss any aspect of this submission in a public meeting or pre-submission consultation context. Respectfully submitted, Walnut Hill Medical Healthcare Reimbursement & Commercialization Strategy Consulting Dallas, TX August 18, 2026 END OF COMMENT
Original source ↗
Source directory

All 32 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Brandon KaplanIndustry · Sep 8, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026OrinyxIndustry · Sep 7, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Shara GospelIndustry · Aug 24, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Joel GrunhutPublic / patients · Sep 7, 2026Xiangyu Guo (Independent Researcher)Public / patients · Sep 13, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026