FDA GenAI discussion / Question 1 of 26

Does a two-axis framework, AI device activity and the consequence of relying on an incorrect output, capture the dimensions of risk?

Full FDA question

Does the two-axis risk framework, organized around device activity and the consequence of relying on an incorrect output, appropriately capture the dimensions most relevant to the risk of a GenAI-enabled software function? If there are additional dimensions—such as the reversibility of a resulting action, the availability of downstream safeguards, the time pressure of the deployment setting, or the traceability of the output (i.e., to primary source materials)—that should be represented in a risk framework, please describe and provide examples of how they should be represented.
Read the FDA discussion paper ↗

34 of 95 submissions reference this question.

All audiences
Alfred McBrideIndustry · Aug 18, 2026AnonymousIndustry · Aug 18, 2026Cara AI (Renee Dua, MD)Industry · Aug 18, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Matthew Collins (Quality and Regulatory Executive)Industry · Sep 15, 2026Nathan SabichIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026QRx PartnersIndustry · Aug 24, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026VivaSecurisIndustry · Aug 25, 2026Vizma CarverIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Mitchell BergerAcademia / other · Aug 25, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026
22 Industry6 Clinicians2 Public / patients4 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/1
Filter by audience
Question 1 · Public feedback

Positions on this question

21 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13.

Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Sam Rosenthal (Red Kit)

Industry · Sep 9, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 — additional dimensions. Two of the dimensions the paper floats are, in my setting, the whole story, and I would ask that they be first-class rather than "additional." Time pressure and the availability of downstream safeguards are the same axis seen from two sides. A patient-facing function used at home with a phone in hand has a downstream safeguard: 911. The same function used by a deckhand on a vessel two days from port, or a miner underground, has none — the app is the last safeguard, and the alternative to the app is not a clinician, it is an untrained person acting from memory or not acting at all. The framework's consequence axis should be read against the care that would actually happen without the device (see Question 15), not against an assumed clinical backstop. In the no-backstop setting the risk of an incorrect output rises, but so does the cost of no output; a framework that only counts the first will push developers to refuse exactly where the user has no one else. Traceability of the output to primary source material is the dimension I would put at the top for lay-rescuer guidance. Bystander first aid is unusual among clinical domains in that the correct action is published, fixed and consensus-based (the AHA and ILCOR resuscitation guidelines and their first-aid counterparts). An output that can be traced step-for-step to a published protocol is auditable in a way an open-ended clinical answer is not: a reviewer can mark each step right or wrong against the source. I would suggest that for domains with a fixed reference protocol, the framework treat protocol fidelity — does the output deviate from the published steps — as the primary risk modifier, and treat generated deviation from protocol as the hazard to be measured, rather than treating all generated text as equally open-ended.
Original source ↗

Navid Farr

Industry · Sep 8, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
R1. Add data exposure as a dimension of risk (Questions 1 and 3) This follows from S1. The two-axis framework measures the consequence of relying on an incorrect output. It does not measure the consequence of the data flow that produces the output. Most GenAI- enabled devices will transmit patient inputs — free-text narratives, images, laboratory values, medication lists — to a hosted foundation model operated by a third party. Whether those inputs are retained, used to train future models, accessible to the model operator's staff, or transferred across jurisdictions is invisible to the patient and, in many cases, to the sponsor. The paper does not use the word "privacy" once. I recommend that CDRH treat data exposure as a risk dimension in its own right, either as a third axis or as a mandatory modifier that can move a function upward on the consequence axis. A patient-facing symptom checker that sends a full clinical narrative to an external model is a different device from one that runs the same model on-premises, even if the outputs are identical. Concretely, the framework should account for: whether patient inputs leave the sponsor's control; whether they are retained beyond the session; whether they may be used for model training or improvement; and whether the patient has been told any of this in plain language. I recognize that data protection is formally the domain of other authorities. But a data-flow failure becomes a device safety failure the moment it changes patient behavior. Patients who do not trust where their information goes withhold it, and a device given incomplete inputs produces worse outputs. The European approach, which treats purpose limitation and data minimization as design requirements rather than as after-the-fact compliance, is a better fit for a technology whose inputs are unbounded. At minimum, CDRH should expect sponsors to document the data flow of a GenAI-enabled function as part of the device description, and should treat a contractual and technical prohibition on training on device inputs as a baseline safeguard for patient-facing functions.
Original source ↗

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Response to Discussion Questions 1, 2, and 14 The two-axis framework (device activity × severity of harm) is sound. Question 1 asks whether additional dimensions — including the time pressure of the deployment setting — should be represented. My answer is yes, and specifically: the degree of human oversight should be treated as an empirical property of the deployment setting, not as a design feature that is present or absent. In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override option does not guarantee the override will be used. Automation bias is well documented, and clinicians under time pressure defer to confident outputs. Appendix A element E.4 recognizes automation bias, but treats it as a communication-quality attribute of the device. I would encourage CDRH to also treat it as a modifier of position on the activity axis: a function nominally placed at “acts with continuous HCP supervision” may in practice operate closer to autonomy if the supervision is not exercised. I encourage FDA to: Treat “degree of human oversight” as a property to be demonstrated in representative use conditions — time-pressured, multi-patient, real interface — rather than asserted in labeling. Ask sponsors to show evidence that intended users can and do detect incorrect outputs in representative workflows. This is an override-rate and detection-rate measurement, and it is precisely the kind of human-AI team evidence contemplated in Question 14. Note that the paper’s own observation — that a “talk to your doctor” statement may not make an output less directive — applies symmetrically, and bears on Question 2: an override interface that is never used provides no oversight. Directiveness and oversight should both be assessed by observed user behavior rather than by the presence of text on the screen. 3. Postmarket monitoring: borrow the re-credentialing model
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

The framework is insufficient

Rejects the two-axis framework as sufficient risk assessment; recommends probability and severity analysis.

Read the source passage
Question 1: The two-axis framework The two-axis framework is not sufficient for risk assessment. Device activity does not equal the probability of harm. A highly automated device can have narrow limits and strong controls. An advisory device can cause harm through automation bias. CDRH should continue to use the standard concepts of probability and severity. The risk analysis should include: The unsafe device behavior. The possible harm. The probability of the harm. The available risk controls. The remaining risk after the controls operate. The analysis must include more than an incorrect output. A GenAI device can also fail through omission, harmful agreement, delayed escalation, or an unauthorized action. Reversibility, time pressure, traceability, and available safeguards should modify the risk analysis. CDRH does not need more graphical axes.
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
3.1 Question 1 - Risk assessment The FDA asks whether a possible two-axis framework based on device activity and the consequence of relying on an incorrect output captures the principal dimensions of GenAI risk, while also raising factors such as reversibility, safeguards, time pressure and traceability (FDA, 2026a). Sitora supports the two-axis concept as a strong foundation but recommends explicit treatment of four additional modifiers: evidence traceability, independent reviewability, reversibility and time to harm. Two systems may provide the same recommendation while presenting different risks if one gives the user a transparent evidence chain and the other provides only an opaque conclusion. Risk modifier Evidence traceability Independent reviewability Reversibility Time to harm Table 1. Proposed modifiers to the FDA risk framework. Source: Sitora analysis based on FDA (2026a).
Original source ↗

Krishna Koka

Academia / other · Sep 1, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 (additional risk dimensions). An implant proposed by a GenAI function sits beyond the top of the activity axis: the output is an object, not advice, and remediation after implantation means revision surgery. Both reversibility and traceability, which Question 1 raises, should be represented. I suggest a physical-output modifier that presumptively places any function generating implant geometry at the high-consequence end regardless of how the output is worded. The paper’s reasoning for measurement functions applies with equal force here: a surgeon can judge anatomic fit visually but cannot see internal porosity, a stress concentration, or an inadequate fatigue margin, so the user cannot independently evaluate the properties that matter most.
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Response to Discussion Question 1 – Risk Assessment I support FDA's objective of establishing a risk-proportionate framework for the regulation of GenAI-enabled medical devices. The proposed two-axis framework based on device activity and the consequences of relying on an incorrect output may provide a useful and pragmatic starting point for high-level regulatory decision-making. However, I believe that the framework would benefit from a clearer definition of the type and purpose of the “risk assessment” being performed. The Discussion Paper appears to focus primarily on an initial assessment used to derive regulatory requirements for a given device. This is an important phase to consider. However, this type of risk assessment may be confused with the product-specific risk assessment that is performed during product development and throughout the product lifecycle. In the Discussion Paper, it seems that the term “risk” is used at both levels even though the underlying concepts do not fully coincide. At the first level, only information that is already defined and can be reliably supported when the initial regulatory criticality is assigned (e.g., in the device description) should be used. Information that becomes available only through subsequent development and risk management should not yet be relied upon. For example, reliable product-specific estimates of the probability of occurrence of harm, or of probabilities associated with specific failure and harm scenarios, may often not yet be available at this stage. This ambiguity becomes even clearer when considering the definition of risk in ISO 14971. In ISO 14971, risk is defined as the “combination of the probability of occurrence of harm and the severity of that harm”. As discussed above, product-specific probabilities may often not yet be sufficiently characterized during the initial phase. Consequently, the term “risk” as used for the initial regulatory assessment may be interpreted differently from the ISO 14971 definition of product risk, which could create ambiguity. In relation to the two-axis framework as proposed in the Discussion Paper, this definition of risk overlaps with the consequences axis. Some passages of the Discussion Paper appear to use the consequences axis more broadly by also considering factors that affect reliance on an incorrect output or the likelihood that such an output results in harm. This seems to be more related to product-specific risk assessment. However, this relationship is not explicitly defined. From my perspective, the Discussion Paper should clarify more explicitly the context in which the term risk is used. It should always be clear at which level the term operates. The two levels should be clearly separated and the terminology should indicate the respective context. I suggest using different terms to avoid confusion: the first level could be described as “criticality” of the product, whereas the second could remain “risk” in accordance with ISO 14971. Furthermore, I suggest separating the two phases and clearly delineating them. I propose the following names for them, while recognizing that these terms are not established regulatory categories. 1. Criticality Assessment / (Regulatory) Criticality Stratification: This phase is used to determine the appropriate level of regulatory oversight and evidence for a defined device and use context. At this stage, only information that is already defined and can be reliably supported should be used, e.g., the type of activity as defined in the Discussion Paper or the severity of a potential harm. Other factors, such as the application context, the required competency profile of users (e.g., lay persons/patients versus healthcare professionals), or other fixed parameters regarding integration of the device into the clinical workflow, may also be considered. This phase could be called Criticality Assessment when it focuses on the general assessment of criticality. Subsequently, I use the term Criticality Stratification because it refers to a categorization into criticality levels that can then be used to determine the applicable regulatory requirements. 2. Product-Specific (Safety) Risk Management This phase refers to the established risk management process according to ISO 14971 that is performed throughout device development and the total product lifecycle, including risk analysis, risk evaluation, risk control, and evaluation of residual risk consistent with ISO 14971. For the following comments, I interpret Section IV of the Discussion Paper primarily as referring to the first phase, i.e., the Criticality Stratification introduced above. The following considerations are based on this interpretation. This initial phase sets the anchor for subsequent regulatory and development steps. Accordingly, only factors that are defined and can be reliably supported at this stage should be used. These may include the following factors where the first two entries primarily correspond
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1: Adequacy of the two-axis framework and additional dimensions. The two axes, device activity and the consequences of relying on an incorrect output, capture the dominant drivers of risk, and the framework is usefully continuous with the benefit-risk logic of existing FDA guidance. I recommend keeping the framework two-dimensional and treating the additional dimensions CDRH lists (reversibility of the resulting action, availability of downstream safeguards, time pressure of the deployment setting, and traceability of outputs to primary sources) as documented, auditable modifiers of a function's position on the consequences axis, rather than as additional axes; more than two axes would compromise usability without adding discriminating power. One further modifier warrants explicit inclusion: the detectability of an incorrect output by the intended user. An error the user can plausibly recognize and discount (for example, a summary contradicting a visible source record) is materially lower risk than an error the user cannot independently verify.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 - Two-axis risk framework and additional dimensions The proposed two-axis framework - device activity on one axis and consequence of relying on an incorrect output on the other - is a useful primary heuristic. It is intuitive, clinically meaningful, and readily scalable from non-directive informational functions to fully autonomous action-taking functions. FDA should retain the two axes rather than converting the framework into an unmanageable multi-dimensional matrix. [1] FDA also appropriately identifies GenAI-enabled in vitro diagnostic, measurement, and signal-processing functions as presenting special risk considerations because their outputs may not be directly assessable by the user, limiting the user's ability to recognize and avoid reliance on an incorrect output. [1] Additional dimensions should function as risk modifiers that influence placement within the framework and the strength of associated controls. Particularly important modifiers include: • Traceability: whether a consequential output can be linked to the patient or encounter, source information or evidence, model/configuration state, and primary materials on which it relied. • Reversibility: whether an incorrect action can be promptly reversed without lasting harm. • Time pressure: whether meaningful human review is feasible before harm can occur. • Downstream safeguards: whether an independent mechanism can constrain, intercept, escalate, or block an unsafe action. • User capability and setting: whether the output is delivered to a patient, generalist, specialist, or another system that may have different ability to independently evaluate it. • Execution authority: whether the function merely provides information, strongly directs action, or can itself cause a consequential action to occur. These modifiers can be operationalized without changing the basic two-axis structure. For example, a high-consequence action-taking function with limited reversibility, little time for human review, and no downstream gate should require materially stronger premarket and postmarket assurance than a reversible, traceable, HCP-supervised function at the same nominal activity level. The amount and type of assurance should also be proportionate to the extent to which clinically relevant system state lies outside the sponsor's direct control. A bounded, recommendation-only function may be adequately addressed through comparatively conventional competency and lifecycle controls, whereas a high-consequence agentic function dependent on external models, tools, clinical data sources, institutional policy, or changing configuration may warrant substantially stronger provenance, configuration- lineage, authorization, and postmarket controls. This is consistent with a least-burdensome approach: the objective is not maximum control for every device, but the minimum 4 information and control functions necessary to address the relevant safety and effectiveness questions for the device's actual risk profile. [15] For purposes of this comment, external dependency refers to a clinically material model, data source, retrieval source, tool, service, policy state, permission state, or infrastructure component whose relevant state or modification is not wholly controlled by the device sponsor.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 — Is the proposed two-axis risk framework sufficient? Risk should be measured on five axes, not two. Two outputs that look equally “wrong” on paper can carry very different risk depending on what the system is authorized to do about it. An AI drafting patient-education text and an AI independently adjusting an insulin dose can both be wrong — only one can act on that error without a human in between. I recommend five risk dimensions: • Consequence — how severe is the harm if the system is wrong? • Clinical agency — how much authority does the system have to influence or initiate action? • Reversibility — can an incorrect action be easily undone? • Time-to-harm — how quickly can a wrong output cause irreversible harm? • Human recoverability — will a qualified human likely catch the error before harm occurs? The governing principle: the more clinical agency a system holds, the stronger the evidence required before that agency is granted.
Original source ↗

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 asks whether additional dimensions, including the traceability of the output, should be represented in the risk framework. I suggest that traceability of the output to source materials and traceability of the action to an accountable person are distinct concerns, and that the second is load-bearing for the activity axis specifically. The activity axis positions a function according to how independently it directs or takes action. The paper states that a function acting with continuous professional supervision is meaningfully different from one acting in fully autonomous fashion. That is correct and it is the right distinction to draw. It also creates an evidentiary obligation: for the distinction to carry regulatory weight across the total product life cycle, a deployed system must be able to demonstrate which side of it a given action fell on. Three identities separate when an agent acts in a clinical system, and each answers a different question: Authenticating identity. Whose credentials were presented to the target system. Acting identity. Which entity performed the action. Authorizing identity. Whose clinical intent the action represents. For direct human action all three coincide, which is why one field sufficed. For an agent operating under a clinician’s authorization they are three different things, and the position of the function on the activity axis is a claim about the third one. A framework dimension for attribution traceability would ask whether a device’s emitted records preserve the distinction between these identities, and would place a device that cannot preserve it at higher risk than one that can, holding activity and consequence constant. The rationale is that supervision claimed at authorization but unverifiable in deployment is not equivalent to supervision that can be shown. A concrete illustration. A publicly documented reference implementation, FHIR Agent Studio, demonstrates the gap without any of the shortcuts that would make it an unfair example. It comprises twelve clinical agents running against a synthetic FHIR repository on a widely deployed health data platform, and its author describes it as a prototype rather than a product. Its treatment of transparency is thorough. Each run retains the assembled prompt and the raw model response, a replayable trace of the full message flow including the record read, the query, the retrieval step, and the model call, and a response history recording each model output with its prompt, a timestamp, and a hash. Every action the agents propose is a draft that a clinician approves before anything is written. That design is careful, and the constraint it runs into is structural rather than a defect of the implementation. The trace is held in the platform’s interoperability layer, not in the standard audit record. It does not travel with the clinical data and it is not what an oversight query reads. At the approval step, what enters the clinical record is a resource attributed to the approving clinician. The agent’s participation persists only in vendor-specific platform telemetry, correlated to the clinical record by nothing durable within it. Reading the record afterward, an investigator cannot distinguish a resource the clinician composed from one an agent drafted and the clinician approved. This is precisely the human-supervised band of the activity axis in Figure 1. A well-built implementation, on a mainstream platform, operating exactly as intended, produces a record from which the supervision relationship cannot be recovered. The retention of prompts and model outputs illustrates a related point: that content is appropriate for a synthetic corpus, but persisting it over real patients creates a disclosure surface subject to minimum-necessary and, for substance use disorder records, 42 CFR Part 2 constraints. Attribution is better carried as scalar metadata in the audit record than as retained model content, because the metadata form can be preserved indefinitely without creating that surface. 3. Response to Question 20: attribution is a precondition for machine-based supervisory agents
Original source ↗

Mitchell Berger

Academia / other · Aug 25, 2026

Keep it, but add or change elements

Keeps the two functional axes inside a proposed three-layer model that adds model and lifecycle risks.

Read the source passage
Pages 4–6 and Discussion Question 1: Limitations of the Two‑Axis Model: The two-axis model discussed by FDA is a good starting point for discussion but does not capture such upstream risks as hallucinations, algorithmic bias and incorrect data. While these axes capture autonomy and harm severity, they do not account for upstream risks such as hallucinations, algorithmic bias, incorrect or contaminated training data, or opaque model provenance. The model is also primarily utilitarian, emphasizing consequences but not reflecting governance priorities, ethical guardrails, or healthcare duties such as privacy, confidentiality, and transparency. In addition, GenAI risks are not static. They may change as models are updated or not updated, as models drift, as data degrades, or as context and use‑cases shift. A static two‑axis model cannot capture 2 Page 1 Fatehi F, Samadbeik M, Kazemi A. What is Digital Health? Review of Definitions. Stud Health Technol Inform. 2020 Nov 23;275:67-71. doi: 10.3233/SHTI200696. Wienert J, Jahnel T, Maaß L What are Digital Public Health Interventions? First Steps Toward a Definition and an Intervention Classification Framework J Med Internet Res 2022;24(6):e31921; Olivia A. Stein, Audrey Prost, Exploring the societal implications of digital mental health technologies: A critical review, SSM - Mental Health, 2024(6): 100373, https://doi.org/10.1016/j.ssmmh.2024.100373 2 https://www.fda.gov/science-research/artificial-intelligence-and-medical-products/fda-digital-health-and-artificial-intelligence-glossary- educational-resource; See e.g., ONC/ASTP’s glossary at https://www.healthit.gov/topic/health-it-and-health-information-exchange- basics/glossary and NIST’s at https://csrc.nist.gov/glossary 3 https://csrc.nist.gov/glossary/term/foundation_model 4 National Academy of Medicine; The Learning Health System Series; Elliott A, Krishnan S, Sarich T, et al., editors. Generative Artificial Intelligence in Health and Medicine: Opportunities and Responsibilities for Transformative Innovation. Washington (DC): National Academies Press (US); 2025 May 16. 3, RISKS OF GENERATIVE ARTIFICIAL INTELLIGENCE IN HEALTH AND MEDICINE. Available from: https://www.ncbi.nlm.nih.gov/books/NBK615587/;https://oecd.ai/en/genai/issues/risks-and-unknowns 5 National Academy of Medicine; The Learning Health System Series; Elliott A, Krishnan S, Sarich T, et al., editors. Generative Artificial Intelligence in Health and Medicine: Opportunities and Responsibilities for Transformative Innovation. 2025. Available from: https://www.ncbi.nlm.nih.gov/books/NBK615587/ these dynamic, lifecycle‑dependent risks, underscoring the need for a more expansive and adaptive regulatory framework. Instead of a two-axis model, I would suggest a three-layer model: Layer 1, Functional: Consequences (same as current graph/model) and Activity (same as current graph/model). Layer 2, Model Risks: addresses hallucinations, biases, data contamination, poor or incorrect training data, susceptibility to adversarial prompts; and vulnerability to hacking, model manipulation, or other security threats. Layer 3, Lifecycle Risks: model drift, degradation, need for updates, stability. Action-taking and Action-directing functions: The paper states that “CDRH is also considering whether a patient facing informational function may not become any less directive because it includes a ‘talk to your doctor’ or an ‘I am not a medical professional’ statement in addition to the ‘action-directing’ information.” The assumption that patients will reliably seek clinician input is unrealistic. The more likely outcome is that patients will simply consult another AI model or tool to cross‑reference the first output. FDA should treat action‑directing outputs as higher‑risk regardless of disclaimers. Patient-facing informational functions versus health care provider-facing informational functions and Discussion Questions 2 and 3: Patient‑Facing Informational Functions: Patient‑facing functions should provide clear, accessible, bottom‑line information that supports autonomy and shared decision‑making. Interfaces may be simplified, graphical, or user‑friendly, but patients should have full optional access to underlying data, uncertainty indicators, and model limitations. Health Provider‑Facing Informational Functions Provider functions may be more complex, detailed, and data‑rich. Clinicians may need intermediate outputs, confidence scores, model rationales, and contextual factors that would not benefit many patients. These functions support clinical judgment, diagnostic reasoning, and safe integration of GenAI outputs into care (page 8). Outputs can be more directive because they are interpreted within the clinician’s professional judgment and scope of practice. FDA can consider the following factors when assessing patient‑facing versus provider‑facing informational functions and the potential risks associated with each. • Wording and tone: For patients, wording should be
Original source ↗

VivaSecuris

Industry · Aug 25, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1. The two-axis framework is useful, but its output should be modified by operational context and demonstrated control effectiveness. Device activity and consequence of incorrect output are appropriate primary axes. They should establish inherent risk. Residual risk, however, cannot be understood without modifiers that determine whether harm can be prevented, detected, interrupted, reversed, or reconstructed. Modifier Regulatory significance Whether the resulting action can be reliably undone Reversibility before harm occurs. Whether there is sufficient time for detection, Time to harm human review, or safe interruption. Whether a qualified human or separately evaluated Independent oversight control can prevent or stop the action. Whether outputs and actions can be attributed to Traceability sources, components, policies, and decision pathways. Whether one output changes future permissions, State persistence memory, treatment state, or downstream behavior. How broadly an error can propagate across patients, Exposure and scale sites, or connected devices. Whether uncertainty is measured, calibrated, Uncertainty visibility communicated, and used to trigger deferral. FDA can preserve the simplicity of the two-axis figure by treating these factors as a structured modifier profile rather than adding many graphical axes. Sponsors should state inherent risk, controls relied upon, evidence of control effectiveness, and resulting residual risk.
Original source ↗

QRx Partners

Industry · Aug 24, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1: Two-Axis Risk Framework The proposed framework based on device activity and the consequence of relying on an incorrect output is useful for initial risk characterization, but it should not substitute for device-level risk analysis. Risk also depends on how an erroneous output could progress to a hazardous situation. Relevant factors may include the user's ability to recognize the error, opportunity for intervention, reversibility, time to respond, and downstream safeguards. FDA should treat the two-axis framework as a screening or organizing tool rather than a classification matrix. Further evaluation should consider the GenAI-enabled function within its intended use environment and overall risk management process.
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 — Is the proposed two-axis risk framework sufficient? FDA asks whether assessing both the function performed by a generative-AI medical device and the consequence of relying upon an incorrect output adequately captures risk, and what additional considerations may be necessary. Response The proposed framework is a useful starting point, but I do not believe that “incorrect output” captures the full universe of foreseeable patient harm. Healthcare AI can produce an entirely correct output and still create an unsafe interaction. Consider a patient who asks a health-benefits or navigation AI: “Can you send me my insurance card?” The AI retrieves the correct insurance card immediately. The task is completed flawlessly. But imagine that the patient is a stressed, obese, middle-aged executive with diabetes sitting alone in his home office late at night. He is sweating, nauseated, and experiencing what he describes to himself as severe heartburn. He is considering seeking care and is trying to find his insurance information first. He never asks the AI: “Am I having a heart attack?” The AI answers the question he asked. He returns to work. At 5:00 a.m., his spouse realizes he never came to bed and finds him dead in his home office. In that scenario, the AI did not hallucinate. It did not give incorrect clinical advice. It did not malfunction. It did exactly what it was asked to do. That is the safety problem. And it illustrates a principle that I believe deserves explicit consideration in the regulation and evaluation of patient-facing AI: AI can be smart without understanding human behavior. A human being may not ask, “Am I having a heart attack?” Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 2 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices He may ask for his insurance card. A patient worried about medication toxicity may ask, “How much does this prescription cost?” A diabetic patient with a potentially limb-threatening foot problem may ask, “Can you find me a podiatrist?” A patient whose symptoms have worsened may ask, “Can I move my appointment to next month?” The literal request may be administrative while the underlying need is clinical. Risk assessment therefore should include not only the consequence of an incorrect output, but also:  the consequence of an incomplete interaction;  failure to identify a latent clinical need;  failure to recognize safety-relevant context;  risk created by omission;  risk of inappropriate reassurance;  risk that administrative or financial information causes delay or abandonment of necessary care;  whether the system has an opportunity and responsibility to ask an appropriate follow-up question;  whether qualified clinical escalation is available;  the time sensitivity and reversibility of the potential harm; and  the cumulative effect of a multi-turn interaction. This concern is not hypothetical. Amazon Science’s 2026 PatientAgentBench evaluated 10 models across four families on 1,200 patient- facing scenarios and found an important divergence between successful task execution and safe clinical behavior. Triage quality was the most discriminating dimension, and the authors reported that agents often acted on administrative requests without clinical screening. [3][4] FDA should review that work as it considers how patient-facing and agentic AI should be evaluated. PatientAgentBench paper: https://arxiv.org/abs/2607.25485 Amazon Science reference implementation: https://github.com/amazon-science/PatientAgentBench The significance of this finding extends beyond the particular agents evaluated. Traditional technology evaluation tends to ask: Did the system complete the requested task correctly? Healthcare safety requires an additional question: Should the system have completed that task without first recognizing that something more important might be happening? PatientAgentBench provides empirical support for incorporating this distinction into FDA’s risk framework. The appropriate regulatory concept is therefore broader than output accuracy. Task completion is not an adequate healthcare-AI safety metric. For patient-facing systems in particular, FDA should evaluate whether the AI can recognize when the patient’s literal request may be only a proxy for an underlying healthcare need.
Original source ↗

Nathan Sabich

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
FDA Question 1: To start, the two-axis framework is a sound starting point but is insufficient for GenAI. It omits three dimensions that are critical for real-world assessment: (a) traceability of the output to source data, (b) the availability and effectiveness of downstream human oversight, and (c) the temporal context of deployment (real-time clinical decision vs. asynchronous review). A GenAI function that generates a patient-facing summary with no clinician intermediary poses fundamentally different risk than one that produces a draft note for clinician review, even if both occupy the same position on the two-axis grid. An Architecture to Address This: I have created an architecture that explicitly structures risk mitigation through multiple dimensions beyond device activity and consequence. The HITL Validation Points across the lifecycle define distinct human oversight checkpoints, each with named roles and defined responsibilities. I utilize a transparency and labeling framework that requires documentation of intended use, intended user, known limitations, subgroup performance, model characteristics, output interpretation guidance, and update/version information. I created a multi-gate quality checkpoint system that provides progressive risk assessment at design, data integrity, model metrics, model acceptance, business review, and final live deployment, each with designated KOL reviewers. Recommendation to FDA: Expand the framework to a four-axis model: (1) device activity (current Axis 1), (2) consequence of incorrect output (current Axis 2), (3) degree of human oversight available at the point of output delivery (ranging from fully autonomous to mandatory clinician review), and (4) traceability of output to verifiable source data (ranging from fully traceable to opaque/generative). This four-axis model would better capture the actual risk surface of GenAI-enabled devices and align with IMDRF's lifecycle management framework, which emphasizes human oversight as distinct a governance dimension
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1 — Dimensions of the risk framework The two axes — device activity and the consequence of relying on an incorrect output — are necessary but not sufficient. Together they describe severity and a proxy for autonomy. They do not describe the probability that an incorrect output is acted upon, which is where generative systems differ most sharply from the software CDRH has regulated to date. This matters because ISO 14971, which manufacturers already apply and document, defines risk as the combination of the probability of occurrence of harm and the severity of that harm. The paper’s framework captures severity well and probability not at all. A device whose errors are obvious on inspection and a device whose errors are fluent, confident, and internally consistent may occupy the same cell of the proposed grid while presenting materially different risk. Fluency is precisely the property generative models optimize for, and it is an anti-correlate of detectability. I recommend CDRH add detectability — the likelihood that an incorrect output is recognized as incorrect before it is relied upon — as an explicit dimension, and treat reversibility, downstream safeguards, time pressure, and output traceability as determinants of detectability rather than as independent axes. This preserves a two-dimensional grid that remains usable in practice while giving the omitted variable a defined home. I further recommend that CDRH state explicitly how this framework relates to a manufacturer’s existing ISO 14971 risk management file. If the answer is that the grid is a communication and triage device that draws on the 14971 file rather than a separate analysis, that should be said. Absent such a statement, manufacturers will reasonably assume they must maintain a second, FDA-specific risk taxonomy alongside the first, which produces documentation burden without a corresponding safety gain. 2 of 19 Docket No. FDA-2026-N-7874
Original source ↗

Cara AI (Renee Dua, MD)

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1. Dimensions missing from the two-axis risk framework The two-axis framework is a reasonable organizing heuristic. We suggest three additional dimensions. Reversibility and time to consequence: a generative output feeding a decision made minutes later carries different risk from one feeding a decision made weeks later after human review. In long-term services and supports, an incorrect activity of daily living score flows into a service authorization a plan reviewer examines, a state may audit, and the member may appeal. The path from output to consequence is slow, documented, and reversible at several points. A framework blind to this dimension treats a home safety observation the same way it treats an emergency department triage output. Depth of downstream human review: the activity axis distinguishes healthcare professional supervision from full autonomy. We suggest a finer distinction within supervision. Oversight ranges from a clinician glancing at a generated summary to a clinician independently reaching a determination and then comparing it against the output. The second is a substantially stronger control and should reduce assessed risk more than the first. Traceability to source evidence: where an output is anchored to a specific artifact the reviewer sees, such as a photograph, a document image, or a recorded statement, the reviewer holds the © 2026 Cara AI, Inc. underlying evidence and evaluates the output directly against it. Traceability of this kind lowers the risk of an undetected incorrect output and belongs in the framework as its own dimension.
Original source ↗

Richard Pescatore, DO (BellyMD)

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Question 1: additional risk dimensions. Two dimensions deserve explicit representation. First, the counterfactual care environment. The consequence of relying on an incorrect output is properly measured against what would have happened without the device, and that baseline varies widely across deployment contexts. A function deployed to a population with ready access to specialty care presents a different risk calculus than the same function deployed to a population whose realistic alternative is no professional input at all. Section V.D.1 and Question 15 touch this in the comparator discussion; it belongs in the risk assessment itself. Second, duration and accumulation of reliance. A longitudinal companion that shapes a patient's understanding of their condition over months is different from a single-encounter tool, because error in the former compounds and becomes belief. Exposure time should modify risk the way dose modifies toxicity.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
FDA Question 1 - Risk framework dimensions Trace ID. TR-Q01 | FDA Q1; Sec. IV.A; App. B; pp. 9-10 / 27-28 BCR response. Use the activity x consequence matrix as a communication layer, not the full closure model. Add a mandatory boundary ledger for reversibility, downstream safeguards, time pressure/time-to-harm, traceability/provenance, user review capability, conversational trajectory, model/dependency change, subgroup performance, tool permissions, and postmarket detectability. Overlay hard gates for irreversible or catastrophic branches. BCR rule basis. BCR-R02,R07,R09,R14,R17 Solution-stack link. S2,S4 Closure evidence. Multidimensional boundary record + scenario tests + hard-gate results Pass / re-open. No essential risk dimension orphaned; all applicable hard gates pass Re-open when: Relevant model/use/tool/workflow change or new harm pathway.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Keep it, but add or change elements

Qualified position: the requested changes or conditions in the passage are part of the position, not treated as unconditional support.

Read the source passage
Response to Question 1: Additional Dimensions of Risk The two-axis framework captures the core of the risk landscape but requires three additional dimensions to be complete. First, third-party model dependency is a material risk factor not reflected in the current framework. A generative AI device built entirely on a proprietary, auditable, manufacturer- controlled model presents a fundamentally different risk profile than a device built on an opaque third-party foundation model over which the manufacturer has no visibility, no contractual governance rights, and no notification of changes. The current framework treats these devices identically. They are not identical. FDA should add third-party model dependency as an explicit risk dimension, with devices heavily dependent on non-disclosed, externally controlled foundation models receiving higher baseline risk classification. Second, reversibility must be treated as a first-order risk factor. An AI- generated recommendation that a clinician reviews and acts upon over the course of hours is materially different from an AI-generated command that triggers an immediate, irreversible action — a drug infusion, a neurostimulation parameter change, a surgical robot maneuver. The current framework's consequence axis captures severity but does not fully capture irreversibility. A separate reversibility dimension would strengthen the framework's clinical fidelity. Third, temporal urgency — whether the device operates in real-time versus asynchronous clinical workflows — creates meaningfully different risk profiles that the current framework does not distinguish. Real-time autonomous action under time pressure (e.g., AI-guided emergency drug dosing) warrants higher scrutiny than asynchronous AI-generated care plan recommendations reviewed by a clinician before acting.
Original source ↗
Source directory

All 34 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026AnonymousIndustry · Aug 18, 2026Cara AI (Renee Dua, MD)Industry · Aug 18, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Matthew Collins (Quality and Regulatory Executive)Industry · Sep 15, 2026Nathan SabichIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026QRx PartnersIndustry · Aug 24, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Richard Pescatore, DO (BellyMD)Industry · Aug 18, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026VivaSecurisIndustry · Aug 25, 2026Vizma CarverIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Shannon KamalakerClinicians · Aug 19, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Krishna KokaAcademia / other · Sep 1, 2026Martin HaimerlAcademia / other · Sep 1, 2026Mitchell BergerAcademia / other · Aug 25, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026