Verida Charter Foundation (Alexander D. Barrett)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ5 · Multi-turn conversations that migrateQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
Comment by Alexander D. Barrett, Senior Fellow, Verida Charter Foundation (ORCID 0009-0003-0359-3163) on "Considerations for the Regulation of Generative AI-Enabled Medical Devices." The full comment is attached.
I research how evidence of a decision travels across organizational boundaries. I am the lead author of an open specification for portable decision evidence: signed records carrying the provenance, authority, and human oversight of a decision to a party that must rely on it without being able to re-derive it. The work began in insurance and regulated trade. Seven of CDRH’s questions turn on the same structure. My research base is weighted toward insurance and trade; I claim no health-sector fieldwork, and nothing in the comment rests on any.
Q1, additional risk dimensions. Reversibility should be represented, but as a gate on Question 18 trade rather than as a third axis. Monitoring detects an error after the action has occurred; where the action is irreversible, detection is not a remedy. Severity does not carry this. Holding activity constant, a severe-but-reversible harm sits at the top of the consequences axis, and a limited-but-irreversible harm at the bottom, so the grid places the case monitoring could remedy where it attracts most premarket scrutiny. The case monitoring cannot remedy where it attracts least. Where the two diverge, severity orders them the wrong way round. A function declaring an irreversible action should be ineligible for reduced premarket evidence regardless of severity.
Q2, the directiveness continuum. An axis defined over the wording of an output can be moved by rewording it without changing anything a user can do, which is why it cannot give manufacturers the predictability CDRH seeks. Defining the axis over the function’s effect on the user’s choice set fixes this: presents options, ranks options, filters options, selects option, executes. That property is determined by design before any output text exists. It also explains, rather than asserts, why a "talk to your doctor" statement changes nothing. The continuum as drawn has no rung for omission, which is a substantive gap: a device that narrows a differential by dropping entries directs action by subtraction, is more directive than one that ranks, and is invisible to any wording-based measure.
Q19, postmarket performance. Periodic re-benchmarking, sample-based clinician review, and degradation monitoring each produce a rate, and none as described requires the population it was computed over to be stated. I offer a finding rather than a recommendation: the specification I work on shipped four released schemas carrying values over populations it never required anyone to state. The defect survived review because a ratio looks complete. The remedy carries no patient data, and the check that reaches silent sampling must sit at the level of the source rather than the record, because records never ingested never enter the total considered.
Q20, machine supervisory agents. The first-order risk is not agent unreliability but that a record proving a review process ran becomes indistinguishable from a record proving a human exercised judgment. Four mechanical properties prevent this: the oversight classification must be re-derivable from the signed record rather than stored as a label; the record must commit by hash to what the reviewer actually saw; it must state whether the reviewer was itself model-assisted; and absence must read as absence rather than as a silent pass.
Q24, third-party model changes. A regime recognizing only vendor-issued notices depends on vendor cooperation, which CDRH itself doubts at Q25. Recording who authored a change notice makes deployer-side detection a first-class evidentiary act and makes vendor silence visible. Manufacturers should demonstrate detection capability rather than only a contractual clause, and resumption of reliance on a withdrawn model version should be fail-closed.
Q25, Foundation Model Master Files. The incentive problem is a disclosure-form problem. A Master File asking for documents will be met with documents written to be safe to hand over. Commitments by hash to specific results let a developer surrender less and commit to more. This is the lesson of the PIP implant case, which is a device lesson: a valid certificate against a quality management system, checked against the paperwork rather than the artifact. Expectations should scale by whether independent re-evaluation of the model is possible at all.
Q26, agentic devices. A checkpoint that could not run must not read as one that passed; undeclared machine involvement must read as undeclared; each action should commit to which input origins were in context; and the length of an action sequence should be committed, since a signed chain can otherwise be truncated and still verify.
I would welcome correction on any point, particularly from reviewers with the clinical grounding I lack.
Attachment
Comment on Docket FDA-2026-N-7874
Re: Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for
Feedback(CDRH Digital Health Center of Excellence, 18 August 2026)
Submitted by: Alexander D. Barrett, Senior Fellow, Verida Charter Foundation ORCID 0009-0003-03593163Date:[DATE OF FILING]
Introduction
I research how evidence of a decision travels across organizational boundaries. I am the lead author of an open,
vendor-neutral specification for portable decision evidence: signed, canonicalized artifacts that carry the
provenance, authority, and human oversight of a decision from the party that made it to a party that must rely on
it without being able to re-derive it. The work began in insurance and regulated trade, where one firm routinely
acts on a decision another firm made. Several of the questions in this discussion paper turn on the same
structure.
I address seven questions: 1, 2, 19, 20, 24, 25, and 26. I have no standing on the clinical, statistical, and programdesign questions and do not address them. The paper expressly permits partial responses.
Two limits should frame everything below. First, my research base is weighted toward insurance and regulated
trade. I claim no health-sector fieldwork and nothing here rests on any. Second, the specification I describe
supplies the evidentiary frame around a result. It runs no evaluations, accredits no evaluators, sets no thresholds,
and contains no clinical content. Three of the ten proposed benchmarking elements, covering clinical knowledge,
information gathering, and quantitative analysis, are outside its reach entirely.
Summary of recommendations
1. Q1.Carry reversibility as a declared property of the function, and make an irreversible action disqualifying
for the premarket-for-postmarket trade contemplated in Question 18, rather than adding a third axis to the
grid.
2. Q2.Define the activity axis over the function's effect on the user's choice set rather than over the wording
of its output. Add a rung for omission, which the current continuum does not represent at all.
3. Q19.Require any postmarket performance figure to be accompanied by a stated population, including a
declaration of whether the sources it was drawn from are wholly accounted for, and a count of any that are
not.
4. Q20.Permit machine supervisory agents only where their output is structurally incapable of being read
downstream as a record of human judgment.
5. Q24.Recognize a deployer-observed change notice as evidence on equal footing with a vendor-issued one,
and require detection capability rather than only a contractual clause.
6. Q25.Structure the voluntary Master File around hash-anchored commitments to specific results rather than
descriptions of process, and scale expectations by whether independent re-evaluation of the model is
possible at all.
7. Q26.For agentic devices, require a checkpoint that cannot be distinguished from one that passed, that
undeclared machine involvement read as undeclared, and that each action commit to its input origins and
to the length of the sequence it belongs to.
Question 1. Additional dimensions for the risk framework
CDRH asks whether additional dimensions should be represented, naming first "the reversibility of a resulting
action," then downstream safeguards, time pressure, and traceability of the output.
Reversibility should be represented, but not as a third axis. It should be a gate on the trade contemplated in
Question 18.
Question 18 asks whether FDA might accept greater premarket uncertainty in exchange for greater reliance on
postmarket monitoring. Monitoring is a detection mechanism. It operates after the action has occurred. Where
the action can be wound back, detection is most of the remedy: the device is corrected, the patient is re-treated,
and the residual harm is the delay. Where the action cannot be wound back, detection is not a remedy at all. It
converts an unknown harm into a known one and leaves the patient exactly where they were.
Severity does not capture this. The grid separates the two cases, but it separates them by severity, and severity
does not govern the Question 18 trade. Holding activity constant, as the paper's own pairing does, the
consequences axis runs from limited through moderate to severe: a severe-but-reversible harm sits at the top of
that axis, and a limited-but-irreversible harm sits at the bottom. Of those two, the function the grid places
highest, and which would therefore attract the most premarket scrutiny, is the one postmarket monitoring could
in fact remedy; the function it places lowest is the one monitoring cannot remedy at all. Severity does not merely
fail to carry reversibility. Where the two diverge, it orders them the wrong way round.
The practical recommendation is therefore narrow. A function should declare a reversibility class for the action it
authorizes, on three ordered values: reversible, where the action can be wound back to its prior state without
residue; partially reversible, where it can be substantially undone but leaves residue; and irreversible. A function
declaring irreversible would be ineligible for reduced premarket evidence regardless of where it sits on the
consequences axis, because the mechanism that would justify the reduction cannot operate on it.
The paper already reaches for this distinction. Proposed benchmarking element A.1 says the element "also
considers ... compliance with human-oversight checkpoints before irreversible or high-consequence actions."
Irreversibility is therefore already doing work in CDRH's own thinking. But it does that work inside a single
illustrative element that would apply to agentic devices alone, rather than in the risk framework that governs
everything. The recommendation is to carry it where it bears weight.
On the other three dimensions, CDRH names them briefly. Downstream safeguards are measurable as a ratio
between the rate at which a system detects and rolls back a harmful output and the rate at which that output
propagates past the point of recall, declared across a stated boundary. Time pressure is a property of the
operating conditions rather than of the output, and is best recorded as its own coarse band, deliberately
independent of the device's confidence in any given result: a device may be well calibrated and still be deployed
into conditions that leave no room to act on its uncertainty. Traceability to primary sources is well served by
recording which method produced an explanation alongside a hash commitment to the explanation itself, so that
a reviewer can later obtain the original and confirm it has not changed, without the explanation needing to travel
with every output.
Question 2. The directiveness continuum
CDRH asks what characteristics of an output could modify risk, and what would give manufacturers "sufficient
clarity and predictability," while recognizing that directiveness is a continuum rather than a binary.
The difficulty CDRH names is structural rather than editorial. The paper observes that directiveness "may depend
on the substance and context of the output, not solely on whether it uses words such as 'recommend,' 'should,'
or 'consider,'" and that a patient-facing output "may not become any less directive because it includes a 'talk to
your doctor' ... statement."
Both observations are consequences of one fact: any axis defined over the wording of an output can be moved by
rewording the output, without changing anything the user can do. That is precisely why such an axis cannot
deliver the predictability CDRH wants. A manufacturer cannot know its classification until the copy is final; the
classification moves when the copy is edited, and a reviewer cannot reproduce the classification without rereading the text.
The axis should instead be defined over the function's effect on the user's choice set. That property is fixed by the
design of the function, is determinable before any output text exists, and does not move when the text is
rewritten. Five ordered values are sufficient:
Value The function Applied to CDRH's lisinopril
examples
Presents options Supplies information; the choice "lisinopril dosages are sometimes
set is unchanged increased when blood pressure
remains above the treatment goal"
Ranks options Weights or orders the choice set; "in situations similar to this,
all options remain visible clinicians often increase the
lisinopril dosage"
Filters options Removes options from view (see below)
Selects option Identifies one option as the one to "I recommend increasing the
take lisinopril dosage"; "increase the
lisinopril from 10 mg to 20 mg
daily"
Executes Performs the action Writing the order
Under this measure, a "talk to your doctor" statement appended to a specific instruction changes nothing,
because the choice set is unchanged. That is the result CDRH says it wants, and it follows from the measure
rather than having to be asserted as an exception.
The continuum as currently drawn has no rung for omission, and this is a substantive gap. A device that narrows a
differential by dropping entries, or that returns three options where five were applicable, is directing action by
subtraction. It is more directive than a device that ranks, because the user can re-weight a ranking they can see
but cannot recover an option they were never shown. No wording-based measure detects this at all: the output
of a filtering device reads as neutral information. This matters most for the measurement and signal-processing
functions the paper separately flags, where the user cannot independently evaluate the basis for the output, and
where a device would be classified as non-directive today.
This also answers Question 5. Because one ladder spans informational and action-taking functions, a multi-turn
device that migrates from non-directive to action-directing traverses rungs on a single axis; its intended use can
be characterized as the highest rung it is designed to reach, rather than requiring a category boundary to be readjudicated mid-conversation.
Question 19. Postmarket performance evaluation
CDRH proposes periodic re-benchmarking, sample-based clinician review, and performance degradation
monitoring.
Each produces a rate. None of the three, as described, requires stating the population over which the rate was
computed. Two rates computed over differently constituted populations are not comparable, and nothing in
either number reveals the difference. A re-benchmarking result cannot be compared to its premarket baseline
unless both declare their denominators.
I offer this as a finding, not a recommendation. The specification I work on shipped four released schemas
carrying values derived from populations it never required anyone to state: an attestation carrying a ratio, two
carrying bands, and one carrying a count. In every case, the numerator traveled, and the denominator did not.
The defect survived review because a ratio looks complete on its face. If it can pass unnoticed in a specification
written specifically to make evidence checkable, it will pass unnoticed in a postmarket monitoring plan.
The remedy is cheap and carries no patient data. Any postmarket performance figure should be accompanied by
a declaration of the population it was computed over, containing: a commitment to the inclusion criterion, the
total considered, the number included, the records excluded with the ground for each exclusion, and, critically, a
declaration of whether the sources the population was drawn from are wholly accounted for, with a count of any
that are not.
Two checks follow, and neither requires additional disclosure. The first is arithmetic: the number included plus
the declared exclusions should equal the total considered, and any shortfall is undeclared exclusion. That defect is
invisible element by element, because every individual figure in a shaped population is well formed; only the
arithmetic exposes it.
The second reaches silent sampling, and it has to sit at the level of the source rather than the record. A
monitoring program that quietly stops ingesting a subset of real-world outputs never brings those records inside
the total it considered, so the arithmetic balances; every element is well formed. The rate stays stable, reassuring,
and entirely misleading. No count computed within the declared bounds will reach it. An operator can state
honestly that a stream exists and is not covered without knowing how many records are in it, which is exactly
what it cannot do record by record. None of this requires anyone to act in bad faith.
Question 20. Machine-based supervisory agents
CDRH asks whether supervisory agents can facilitate postmarket monitoring, and what considerations apply to
"the evaluation and reliability of the supervisory agent itself."
The reliability of the agent is the second-order question. The first-order risk is that a record proving a review
process ran is indistinguishable, downstream, from a record proving a human exercised judgment. Once a
supervisory agent emits monitoring records at volume, every party reading those records later, including FDA, is
reading process receipts. Nothing in their form says whether anyone was positioned to disagree with the device.
This failure mode does not require anyone to act deceptively. It is what happens by default when the artifact
does not carry the distinction.
If supervisory agents are permitted, I recommend that their output be required to be structurally incapable of
being read as human judgment. Four properties achieve this, and all four are mechanical:
8. The oversight classification must be re-derivable from the signed record, never a stored label. A stored label
is exactly how a process receipt comes to read as judgment, and relabelling requires no intent.
9. The record must commit to what the reviewer actually saw, by hash, so that substituting the inputs after
review or replaying an approval onto a different decision is detectable by a party with no access to either
system.
10. The record must state whether the reviewer was itself model-assisted. This is the Question 20 case, and it is
the field that distinguishes a supervisory agent's output from a clinician's at a glance.
11. Absence must read as absence. Where no oversight evidence is present, the correct reading is "not
attested," never a silent pass. Operator-self-signed oversight should be capped below the highest
classification, because the operator is not independent of the thing being overseen.
The same logic bears on sample-based clinician review. A human in the loop who lacks the context, authority,
evidence, or time to challenge an output has not preserved judgment. The institution has relocated liability.
Telling the two apart requires recording not only what the reviewer did, but whether the reviewer was positioned
to do it.
Question 24. Third-party foundation model changes
CDRH asks how a manufacturer can detect, evaluate, and respond to changes initiated by a third-party model
developer, and what mechanisms could provide reasonable assurance.
A change-notice regime that recognizes only vendor-issued notices depends on vendor cooperation, which CDRH
itself doubts in Question 25. The mechanism should therefore record who authored the notice. A record stating
that the deployer observed a change and raised it makes deployer-side detection a first-class evidentiary act
rather than an informal one, and it makes vendor silence visible as a pattern of deployer-authored notices rather
than as an absence of evidence.
A workable notice carries the vendor, the model family, the version before and after, the class of change (update,
withdrawal, subprocessor change, scope expansion), whether the deployer treats it as warranting governance
review, when it takes effect, and a hash commitment to the vendor's own change documentation where any
exists.
Two further recommendations follow.
Require demonstrated detection capability, not only a contractual clause. A contractual notification obligation is
unenforceable given the speed of model changes and provides no evidence when it fails. A manufacturer should
be able to show how it would notice an unannounced change, which in practice means periodic re-benchmarking
against a fixed set with the model version recorded at each run.
Resumption of reliance should be fail-closed. Under the condition Question 23 describes, where the nature and
scope of future modifications cannot be fully prespecified, the safe default is not a broader prespecification. It is
that reliance on a withdrawn or suspect model version may not resume until an affirmative artifact exists
recording what was re-established and by whom. Absence of that artifact should mean the version remains
withdrawn. This inverts the usual default, under which reliance resumes quietly because nothing stopped it.
Question 25. Voluntary Foundation Model Master Files
CDRH asks what would make a voluntary Master File sufficiently useful given that "model developers may have
limited incentive to disclose safety-relevant information."
The incentive problem is a disclosure-form problem. A developer's reluctance is usually not about the fact
becoming known. It is about the artifact leaving their control: the evaluation transcripts, the per-item outcomes,
the internal red-team findings. A Master File that asks for documents will be met with documents written to be
safe to hand over.
A Master File built on hash commitments asks for something different. A developer commits to a specific result
by publishing a digest of it while retaining the underlying record. The commitment is unrepudiable and dated. If
FDA later needs the record, the commitment identifies exactly which record to request and proves it has not
changed in the interim. The developer surrenders far less and commits to considerably more, which is the only
combination that makes voluntary participation rational.
This is the lesson of the PIP implant case, and it is a device lesson. A valid certificate was issued against the
manufacturer's quality management system, not against the implants. Between 1998 and 2008, the notified body
made eight visits to the manufacturer's premises, each announced in advance, and never inspected the business
records or ordered the devices inspected. Silicone that was not an approved material reached hundreds of
thousands of women before the French authority established what had been used and prohibited the implants'
marketing, sale and use in 2010. The failure was not an absent seal. A seal existed, was valid, and was worthless,
because it certified the paperwork rather than the artifact and could not be checked against the thing it vouched
for. A Master File that collects descriptions of a developer's evaluation process repeats that structure. A Master
File that collects commitments to specific evaluation results binds to the artifact.
Two design points follow.
Scale expectations by whether independent re-evaluation is possible. A model distributed as open weights, or on
premises, can be re-tested by any qualified party, and a Master File adds comparatively little. A model available
only to parties the provider admits cannot be independently re-tested by anyone else, making a providersupplied or provider-commissioned attestation the only reliance artifact available. That is the gap the program
exists to close. A Master File program that does not distinguish these cases will over-collect from the developers
who need it least and under-collect from those who are the actual problem.
Make the prespecification CDRH already envisages verifiable after the fact. Section V.B.2 already envisages that
sponsors would prespecify and justify acceptance criteria, and that test methods and acceptance criteria would
be prespecified before testing. I am not proposing that principle; I am proposing a way to check it. Nothing in the
paper gives a reviewer any means of confirming it was honored, and a principle no one can check is one a sponsor
under pressure can quietly decline to follow—section VII. A does not carry the principle across to the benchmark
results a Master File would hold at all. Where a Master File or a submission carries benchmark results, the pass
and fail bounds should therefore be committed by hash at a recorded instant, ideally witnessed independently of
the submitting party's own clock. This does not make a sponsor-developed benchmark independent. It makes one
specific and otherwise invisible abuse falsifiable: moving the acceptance bar after seeing the score. That abuse is
the one most available to a sponsor benchmarking its own device, and a reviewer comparing the anchoring
instant against the evaluation period can detect it in seconds.
Question 26. Agentic devices
CDRH asks what additional considerations apply to agentic devices, and how the elevated risk of autonomous
multi-step action, tool use, and reduced opportunity for human review should be reflected in acceptance criteria.
Four properties are worth requiring of the record an agentic device leaves behind. Each addresses a way an
agentic run can look compliant while not being so, and none requires disclosing the content of the run.
12. A checkpoint that could not run must not read as one that passed. Where an oversight checkpoint was
skipped, timed out, or was unreachable, the record should say so in terms distinguishable from a
checkpoint that ran and approved. Absent that distinction, degraded operation is indistinguishable from
clean operation in exactly the conditions where it matters.
13. Undeclared machine involvement must read as undeclared, not as human. Where a step's actor is not
stated, the correct reading is that it is unknown. A record that omits the actor should not be readable as a
record of human action, which is the default a reviewer will otherwise supply.
14. Each action should commit to which input origins were in context. Prompt injection reaching an agent
through retrieved content or a tool result is not visible in the action itself. Committing to the origins in
context at the moment of the action later distinguishes a device that was manipulated from one that
malfunctioned.
15. The length of an action sequence should be committed. A signed chain of steps can be truncated after the
fact, and a truncated chain can still verify perfectly. Committing to the expected length is what makes a cut
chain detectable rather than merely shorter.
The first and second are the same structural point as Question 20: the failure is that an absence reads as a
positive. The third and fourth are specific to multi-step autonomy, where the record is assembled over time and
can be shaped by what it leaves out.
What I am and am not claiming.
The specification described above is licensed under Apache 2.0 for its schemas and reference code, and CC BY 4.0
for its prose. Its artifacts are canonicalized under RFC 8785, hashed with SHA-256, and signed with Ed25519, and
it is hash-only by construction: commitments travel, content does not. The license terms above describe the
terms the specification carries; I am not claiming here that it is publicly retrievable today.
It is not a broadly adopted public standard. There is no operational certification registry. The artifacts relevant to
this comment differ in maturity, and each one's maturity is recorded and mechanically computed rather than
asserted. I offer the vocabularies here as worked prior art demonstrating that the distinctions CDRH is
contemplating can be made machine-checkable, not as a proposal that FDA adopt this particular specification.
Disclosure. This research is self-funded and received no external funding. To enable empirical validation, I built
prototype implementations of the concepts described here; these are my personal intellectual property and may
later be operated by a successor commercial entity. The specification itself remains openly licensed, so that any
party holding a copy may implement it independently. A disclosure to the same effect appears in my June 2026
research paper.
Availability
A full crosswalk of all ten proposed benchmarking elements and all 26 discussion questions against specific
artifacts, including the elements for which I have nothing to offer, is complete and available to CDRH staff on
request. I would welcome correction on any point above, particularly from reviewers with the clinical grounding I
lack.
Respectfully submitted,
Alexander D. Barrett: Senior Fellow, Verida Charter Foundation ORCID 0009-0003-0359-3163
adb@veridafoundation.org