Nieleze
FDA questions it names
Q10 · Benchmark contamination and saturationQ13 · Synthetic dataQ16 · Independent third partiesQ19 · Postmarket performance evaluationQ22 · Re-benchmarking after a modificationQ26 · Agentic devices
The comment as filed
Please see the attached comment from Nieleze on evidence independence in the evaluation of GenAI-enabled medical devices, addressing principally Questions 10, 13, 16, 19, 22, and 26.
Attachment
Re: FDA-2026-N-7874 — Evidence independence in the evaluation of GenAI-enabled
medical devices
I am a physician-engineer working across biomedical engineering, clinical systems,
simulation, and organizational resilience. I developed the Reliance Lens, a comparative
instrument for examining whether technically different approaches to the same function
actually produce independent operational exposure.
This comment responds principally to Questions 10, 13, 16, 19, 22, and 26.
The Reliance Lens uses the sequence:
Approach → Function → Reliance → Coupling → Independence
Its central proposition is:
Different technologies do not necessarily mean different exposure.
I believe this offers a useful additional question for FDA’s proposed framework:
Does accumulating evidence necessarily produce accumulating
independent information about failure?
FDA’s discussion paper and the comments responding to it describe increasingly rigorous
layers of assurance: dataset documentation and provenance, benchmarking and
construct-validity assessment, clinical confirmation, independent third-party evaluation,
re-benchmarking, postmarket monitoring, clinician review, and supervisory systems.
These mechanisms address different and important problems. The Reliance Lens suggests
examining them collectively as an evidence architecture.
For each evidence-producing approach, the relevant questions are:
Function — what is this mechanism intended to establish?
Reliance — what does it require to perform that function?
Coupling — where do those reliance structures converge with those of other evidence
mechanisms?
Independence — what evidence demonstrates that the resulting exposures are
sufficiently separated?
This distinction matters because institutional independence does not necessarily
establish operational independence.
A third-party evaluator may be organizationally independent while relying on the same
clinical representation, reference standard, data-generating assumptions, model family,
workflow, or upstream signals as the system being evaluated. A dataset may be fully
documented and versioned while remaining a common dependency across benchmarking,
adjudication, and re-benchmarking. A supervisory system may be separately administered
while inheriting the same upstream exposure as the system it supervises. In that case:
device ← shared reliance → supervisory mechanism
may increase the amount of monitoring without creating a commensurate increase in
independent protection.
A concrete example: the re-benchmarking loop
If benchmark evaluation, clinical adjudication, and postmarket re-benchmarking all
depend on the same clinical representation or reference standard, a persistent failure in
that representation can survive model updates and repeated evaluation. The system may
therefore demonstrate stable performance while remaining blind to the same condition.
In simplified form:
shared reliance → missed condition → repeated measurement → apparent stability
This is not a claim that shared reliance necessarily produces correlated failure. It identifies
a testable independence question: whether the evidence mechanisms would behave
differently when the shared reliance is stressed, changed, disrupted, or made unavailable.
The issue therefore differs from ordinary benchmark contamination or model drift. It
concerns whether the evidence-producing system itself has sufficient independence
to discover the failures it is intended to detect.
Evidence volume versus evidence independence
The broader implication is that accumulating evidence does not necessarily mean
accumulating independent information about failure.
A safety case can become more documented, more reproducible, more independently
administered, and more continuously monitored while remaining materially dependent on
the same underlying representation, reference standard, model family, workflow, or
upstream signal.
This suggests distinguishing:
Institutional independence — who performs the evaluation.
Methodological independence — how the evaluation is performed.
Operational independence — whether the evaluator or safeguard depends on materially
different conditions from the system being evaluated.
The first two do not necessarily establish the third.
Possible regulatory extension
For higher-risk GenAI-enabled medical devices, FDA could consider an evidenceindependence assessment alongside existing benchmarking, clinical confirmation, thirdparty evaluation, postmarket monitoring, and supervisory controls.
Such an assessment could identify:
1. the principal evidence and safeguard approaches;
2. the material reliance of each;
3. convergence among those reliance structures;
4. conditions under which those convergences could become consequential; and
5. evidence that apparent redundancy remains meaningful under those conditions.
This would complement, rather than replace, the important measures already proposed in
the discussion paper and public comments.
The central question is:
When multiple sources of assurance appear independent, what evidence
demonstrates that they actually provide independent information about
failure?
The Reliance Lens is designed to make that question explicit and testable. Its outputs are
structured hypotheses about operational reliance and coupling, not claims of causality or
independence without supporting evidence.
Grace Thiong’o
Founder, Nieleze
https://www.nieleze.com
https://gmuthoni.github.io/reliance-lens/