← All 111 filings

Infiligence Inc

IndustryStartupFiled September 22, 20261,700 words · 1 attachmentFDA-2026-N-7874-0114

Themes it raises

6 of the 21 themes in the docket, each with the passage we counted, verbatim.
What makes a function high riskFDA Q1, Q2, Q5
“A multi-turn interaction can begin as informational and become action-directing as new facts, contradictions, symptoms, or user behavior emerge.”
Judging devices the way clinicians are credentialedFDA Q7, Q8
“The objective is not to establish that an AI system is generally “competent,” but that it is competent to perform a defined task, for a defined population, under defined conditions, with a defined degree of autonomy and supervision.”
Whether benchmark results prove anythingFDA Q9, Q10, Q16
“Publicly available benchmarking can provide substantial value even though public benchmarks alone should not establish safety and effectiveness.”
Devices that plan and take actionsFDA Q26
“For agentic systems, safe recovery after tool failure, erroneous tool output, or unsafe action planning should be a gating competency rather than an averaged performance dimension.”
Controlling a device that keeps changingFDA Q22, Q23, Q24, Q25
“When a component changes, the sponsor should perform a Competency Impact Analysis: which competencies could this change reasonably affect?”
Watching the device after it shipsFDA Q19, Q20
“Each important competency should therefore have predefined risk-based requalification triggers or intervals.”
Machine-assisted draft, pending human review. The source text and highlighted passages appear below. Read the filing on regulations.gov ↗

The comment as filed

Comment submitted on regulations.gov. Passages we counted are highlighted.

A GenAI-enabled medical device should not be evaluated as a static piece of software. It should be evaluated as a bounded clinical capability operating within a defined scope, under specified conditions, with measurable competencies, failure tolerances, supervision requirements, and continuous requalification.

Attachment

Attachment, text extracted from the filed document. Passages we counted are highlighted.

Comments on a Competency-Based Approach
for Premarket Evaluation of GenAI-Enabled Medical Devices
Response to FDA CDRH Discussion Paper • August 2026
Submitted by Infiligence Inc. | September 2026

Executive Perspective
We support CDRH’s exploration of a competency-based approach for GenAI-enabled medical devices. Conventional
validation approaches built around bounded inputs, deterministic outputs, and exhaustive test cases are difficult to
apply to systems capable of open-ended interaction, variable outputs, multi-step reasoning, tool use, and evolving
behavior.
CDRH’s proposed combination of device benchmarking and clinical confirmation, scaled to intended use and risk,
provides a strong foundation. We recommend extending that foundation into an operational lifecycle in which
competence is bounded, stress-tested, clinically confirmed, publicly profiled where appropriate, monitored, and
requalified as the device or its operating environment changes.

A GenAI-enabled medical device should demonstrate competence not merely in producing
correct outputs, but in operating safely within a defined competency envelope—including
recognizing the limits of its competence, failing safely, escalating appropriately, and
maintaining those competencies throughout its lifecycle.

1. Define a Competency Envelope
For each GenAI-enabled device, the sponsor should define a Competency Envelope describing the boundaries within
which the device has demonstrated acceptable performance. The objective is not to establish that an AI system is
generally “competent,” but that it is competent to perform a defined task, for a defined population, under defined
conditions, with a defined degree of autonomy and supervision.

• Scope: intended clinical tasks, patient populations, clinical settings, and excluded uses.
• Operating boundaries: autonomy level, permitted tools/actions, environmental assumptions, and required human
oversight.
• Safety boundaries: uncertainty limits, escalation/deferral conditions, known failure modes, and conditions in which
performance has not been established.
This also reinforces an important principle in the discussion paper: the relevant evaluation object is the final user-facing
device in its deployed or representative configuration—not the foundation model in isolation. The practical system is
the foundation model plus prompts, retrieval, knowledge sources, tools, orchestration, guardrails, user interface,
workflow, and human oversight.

2. Make Safe Failure and Recovery an Explicit Competency
CDRH’s proposed safety elements appropriately address safety-critical recognition, escalation, boundary adherence,
uncertainty communication, and clinical deferral. We recommend explicitly evaluating failure competence. For GenAI,
knowing when not to answer or act can be as important as producing the correct answer.

Recognize → Refuse → Escalate → Recover
A competent device should recognize insufficient or contradictory information, refuse actions outside its authorized
scope, escalate appropriately to a human, and recover safely after erroneous information, failed tools, or an earlier
incorrect step. For agentic systems, safe recovery after tool failure, erroneous tool output, or unsafe action planning
should be a gating competency rather than an averaged performance dimension.

3. Evaluate Clinical Trajectories and Stress Conditions
The fundamental unit of GenAI evaluation should increasingly move from the isolated prompt-response pair to the
clinical interaction trajectory. A multi-turn interaction can begin as informational and become action-directing as new
facts, contradictions, symptoms, or user behavior emerge.

Infiligence Inc. • Comments on GenAI-Enabled Medical Device Competency Evaluation
We recommend Trajectory Competency Testing in which scenarios progressively introduce incomplete information,
additional history, contradictory evidence, changing symptoms, patient resistance, cross-domain information, and
potential escalation. A system that behaves safely for ten turns but fails catastrophically on the eleventh should not
appear competent because each response was scored independently.
Trajectory testing should be paired with Competency Stress Testing across four conditions: Normal → Edge →
Adversarial → Failure. This would extend CDRH’s treatment of prompt injection, adversarial prompting, longconversation degradation, contradictory inputs, and information-order sensitivity into a repeatable evaluation method.

4. Establish Minimum Safety Competency Floors
We recommend caution against collapsing all competency measurements into a single aggregate score. Some
competencies should be non-compensable gating criteria. Superior communication or broad clinical knowledge should
not compensate for unacceptable emergency escalation, unsafe dosing behavior, or failure to maintain scope
boundaries. Acceptance criteria could therefore combine overall performance thresholds with minimum floors for
safety-critical capabilities.

5. Introduce Competency Dependency Mapping
Complex GenAI devices increasingly perform tasks through chains of dependent capabilities. A medication
recommendation, for example, may depend on patient-data retrieval, medication normalization, contraindication
detection, dose calculation, recommendation generation, and clinician authorization.
We recommend a Competency Dependency Graph that identifies the capabilities on which each safety-critical device
function depends. Sponsors could evaluate individual nodes where appropriate while also testing the integrated
workflow. This is particularly useful for agentic architectures, where an incorrect outcome may originate from
reasoning, retrieval, tool selection, tool execution, orchestration, or interpretation of a tool response.

6. Treat Benchmarking as Shared Public-Health Infrastructure
Publicly available benchmarking can provide substantial value even though public benchmarks alone should not
establish safety and effectiveness.
CDRH correctly notes both the usefulness of published, reusable tests and the risks
of data contamination, saturation, and limited real-world representativeness. We recommend a hybrid evidence model:

Public benchmarks + Sequestered benchmarks + Sponsor-specific benchmarks + Clinical
confirmation
Public benchmarks can create a common vocabulary through which regulators, clinicians, professional societies,
researchers, developers, healthcare institutions, and patients understand demonstrated capabilities and limitations.
Sequestered and sponsor-specific benchmarks can preserve independence, reduce gaming, and test intended-usespecific risks. Clinical confirmation then establishes whether benchmarked competencies translate into clinically
representative performance.

7. Create a Clinical AI Competency Profile (CACP)
We propose a standardized, versioned, multidimensional Clinical AI Competency Profile (CACP) for a GenAI-enabled
medical device. The CACP would function as a capability disclosure or “nutrition label” for clinical AI—not as an FDA
authorization, a single score, or a leaderboard.

Competency domain Profile status What should be disclosed

Safety-critical recognition & escalation Tested / Not tested Scope, scenario mix, thresholds, material failure modes

Scope & boundary adherence Tested / Not tested Out-of-scope behavior, refusal and redirection

Uncertainty & clinical deferral Tested / Not tested Calibration method, deferral conditions

Clinical domain, population, comparator and acceptance
Clinical proficiency Tested / Not tested
criteria

Infiligence Inc. • Comments on GenAI-Enabled Medical Device Competency Evaluation
Robustness & subgroup performance Tested / Not tested Runtime variation, subgroup coverage and limitations

Agentic / tool-use competency Where applicable Tools/actions, oversight checkpoints, recovery behavior

A public CACP should disclose the benchmark version, device/model version, intended use, evaluation date,
population and setting, relevant uncertainty, and material limitations. The emphasis should be on the conditions under
which competence was demonstrated—not simply a headline number.

8. Separate Foundation Model Capability from Device Competency
We recommend maintaining two related but distinct profiles: a Foundation Model Capability Profile describing general
healthcare-relevant capabilities and limitations of the underlying model, and a GenAI Medical Device Competency
Profile describing performance of the complete device configuration for its specific intended use.

A highly capable foundation model does not automatically create a competent medical
device. Conversely, a carefully constrained device built on a general-purpose model may
demonstrate strong competency within a narrow intended-use envelope.
This distinction complements CDRH’s concept of voluntary Foundation Model Device Master Files while preserving
sponsor responsibility for demonstrating the safety and effectiveness of the actual device configuration.

9. Enable Public, Social, and Industry Profiling—Without a Simple Leaderboard
We encourage an ecosystem in which academia, professional societies, healthcare systems, independent laboratories,
standards organizations, and other qualified parties can continuously evaluate GenAI systems against standardized
competencies. This creates an additional layer of observational evidence and industry learning.
However, social or industry profiling should avoid a single “best medical AI” score. A multidimensional profile is more
informative and less likely to imply interchangeability between systems with different intended uses. Public results
should be versioned and contextualized, and material discrepancies between sponsor-reported, independent, and realworld results should be visible rather than averaged away.
A possible ecosystem is: FDA-defined competency taxonomy → domain-specific definitions from professional
societies/standards bodies → independent public and sequestered evaluation assets → sponsor/device testing →
standardized competency profiles → real-world evidence informing benchmark evolution.

10. Make Benchmark Suites Executable Regulatory Assets
We recommend moving beyond static validation reports toward machine-executable competency evidence. A
benchmark definition could contain the scenario, expected behavioral constraints, scoring rubric, acceptance criteria,
safety threshold, and traceability to the identified risk. This creates the possibility of Validation-as-Code for GenAIenabled medical devices.
The same controlled suite could support premarket qualification, model upgrades, prompt changes, retrieval changes,
tool/orchestration changes, and periodic postmarket requalification. This would make CDRH’s concept of rebenchmarking operational and scalable.

11. Establish a Clinical AI Configuration Baseline and Competency Delta
At authorization, the validated system could have a defined Clinical AI Configuration Baseline covering safety-relevant
versions of the model, prompts, retrieval architecture, knowledge sources, tools, orchestration, guardrails, and userfacing behavior. When a component changes, the sponsor should perform a Competency Impact Analysis: which
competencies could this change reasonably affect?

The resulting Competency Delta Testing could make requalification proportional to the expected impact of the change.
A user-interface change may principally affect communication; a retrieval-source change may affect knowledge and
reasoning; a foundation-model upgrade may affect nearly every competency.

12. Treat Competency as a Lifecycle Property
Competency should not be viewed as something established once at market authorization. Even when the device has
not intentionally changed, clinical guidelines, patient populations, retrieval sources, tools, APIs, foundation models, and
user behavior can change. Each important competency should therefore have predefined risk-based requalification
triggers or intervals.

Infiligence Inc. • Comments on GenAI-Enabled Medical Device Competency Evaluation
Define → Decompose → Benchmark → Stress → Clinically Confirm → Baseline →
Monitor → Requalify

Conclusion
CDRH’s competency-based concept provides an important opportunity to move GenAI medical-device evaluation
beyond deterministic software validation. We recommend extending the framework around four principles: competency
must be bounded; safety includes knowing when not to act; competency evidence should be executable and
longitudinal; and carefully designed public, independent, and sequestered benchmarks should make demonstrated
competency observable across the ecosystem.
The resulting regulatory question becomes more clinically meaningful than “Has this AI model been validated?” It
becomes:

What has this GenAI-enabled device demonstrated that it can safely do, for whom, under
what conditions, with what degree of autonomy and does it continue to demonstrate those
competencies today?

Infiligence Inc. • Comments on GenAI-Enabled Medical Device Competency Evaluation