Specifies safeguards without endorsing the approach
Proposes requirements for using this mechanism, without an explicit overall endorsement or rejection. Kept separate from support.
Read the source passage
5. Question 20: Supervisory agents Manufacturers should define the supervisor's assigned role, the evidence it can access, and its intervention authority. Retrospective screening, triage for human review, and blocking an action require different evaluations. The manufacturer should demonstrate the protection claimed for the selected role. Brandon Kaplan | Individual capacity Page 4 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Evaluators should use task-appropriate reference criteria with qualified clinical or technical adjudication. They should separate evaluation cases from development and tuning cases, document reviewer independence and conflicts, and define how to resolve disagreement. They should use evidence beyond the primary agent's explanation or the supervisor's agreement to establish correctness. Manufacturers should report missed failures by type, false alarms, response times, and uncertainty. They should test shared failure modes and measure the supervisor's added contribution through a justified comparison, such as matched trials with and without its protection in a safe test environment. Choosing different vendors does not by itself establish independence. Two models receiving the same outdated record could agree while missing the selection error; evaluators should check whether the supervisor can access the evidence needed to detect it. The evaluation should include missed unsafe actions, inappropriate blocking of necessary actions, and human-review workload. Manufacturers should use the sampling safeguards in Section 2 to look for errors in unflagged cases. They should assess changes to the supervisor under the same impact-based process used for the device. Manufacturers should define and test the response to an unavailable, delayed, or inconclusive supervisor. They should prevent an uncontrolled bypass of a supervisor needed for safe operation. The evidence burden should follow the claimed protection; this recommendation does not require human review of every output. 6. Investigation records and documentation FDA's software-submission guidance addresses traceability and risk-management documentation, and its cybersecurity guidance addresses security controls and testing. [2, Section VI.C; 5, Section V] I recommend that manufacturers extend those processes to cover agent execution and reuse relevant evidence. They should connect each critical safety claim to its hazard, control, evaluation, and remaining limitation. Manufacturers should retain records needed to investigate foreseeable failures, including configuration identifiers, relevant inputs and source versions, user-facing outputs, approval events, tool requests and results, and resulting state changes. They should correlate records across the agent and connected systems, with enough timing information to reconstruct the sequence. Investigators should distinguish proposed, attempted, blocked, completed, and uncertain actions, and confirm consequential effects against the affected system's records. Manufacturers should protect those records from unauthorized alteration or deletion, including by the agent. They should detect loss of required records and respond according to its safety impact. An interruption in retrospective logging may permit continued operation under a documented limit; loss of evidence needed for an immediate safety decision may require restricting the affected function. Manufacturers should test both conditions. Manufacturers should justify the patient information retained, its purpose, retention period, and access controls. They may use protected references if authorized investigators can retrieve the relevant historical versions for the required retention period. A current chart alone cannot establish what the device saw earlier. Manufacturers should document gaps and protect any additional content needed for investigation under applicable privacy and recordkeeping obligations. Investigators should rely on observable actions and state changes to establish what occurred. A model-generated explanation presented to a user may itself be relevant evidence, but it does not prove that the described actions occurred. This recommendation does not require disclosure or retention of internal model reasoning. Brandon Kaplan | Individual capacity Page 5 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Manufacturers should rehearse investigations using the records and access arrangements available after deployment. Reviewers should be able to identify the configuration, reconstruct consequential actions, and state unresolved gaps. Manufacturers should preserve the exact content shown to users when its wording matters to safety. Investigators need not reproduce that content by running the model again. Conclusion I recommend that FDA ask manufacturers to support critical safety claims with tested controls and documented limitations. Monitoring and response plans should identify who can intervene, what they can do, and whether they can act within th
Original source ↗