Synapstak Ltd (Nicholas P M Baker)
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ5 · Multi-turn conversations that migrateQ7 · The competency-based approachQ9 · The benchmarking structureQ18 · Trading premarket certainty for postmarket monitoringQ19 · Postmarket performance evaluationQ20 · Machine-based supervisory agentsQ21 · Clinicians, institutions and societiesQ22 · Re-benchmarking after a modificationQ23 · PCCPs for GenAI devicesQ24 · Third-party foundation model changesQ25 · Foundation Model Master FilesQ26 · Agentic devices
The comment as filed
Submitted on behalf of Synapstak Ltd (United Kingdom). Full comment attached.
Summary. The two-axis framework is a sound heuristic, but the activity axis is assessed by assertion, from intended use and the substance of outputs. A more auditable basis is what the generative component is architecturally permitted to write and to initiate. A component that cannot write to the record of fact, cannot close its own predictions, cannot initiate an irreversible action except through a deterministic gate, and cannot weigh a stop command is held at a declared activity level by construction, whatever its outputs say and however a conversation drifts. I call this arrangement a write-authority partition and define it by seven verifiable properties in the attachment. In ISO 14971:2019 terms (clause 7.1, Annex A.2.7.1) it is inherently safe design; most current GenAI guardrails are protective measures or information for safety.
The comment is written to be read with others on this docket: 0052 (the test a safeguard must meet to count at stratification), 0054 (separating evidence from inference), 0100 (a deployer’s request that architectural controls count as evidence), 0061 (field observations of self-authored device records), and 0076 and 0094 (opposing positions on postmarket reliance).
Recommendations (Questions 1, 2, 5, 7, 9, 18–26):
Treat the generative component’s write authority as the basis on which activity-axis placement is claimed and verified.
In S.3, distinguish computed from generated confidence and treat generated confidence presented as computed as a safety failure; add a Record Integrity benchmarking element verifying no write path from the component to the record of fact, that unmet predictions surface as failures, and that monitors are passive.
Bind competency evidence to a build fingerprint over the whole deployed configuration; fingerprint change is the re-benchmarking trigger, including for third-party model changes.
Postmarket reliance is defensible where the monitored record is honest by construction, and not otherwise; supervisory agents must satisfy the same partition and pass a passivity test.
For agentic devices: gates on irreversible actions act on witnessed state, not the agent’s account; acts are recorded with their consequences; stops are reflexes outside deliberation; the audit log is one the agent cannot write.
Add a write-authority declaration to the model card, re-verified whenever the fingerprint changes.
Limits are stated in the attachment: the partition protects the record, not the world; it does not reduce confabulation; it does not protect inputs to the deterministic layer; and it does not keep design documentation in step with the build.
Attachment
Comment on Docket No. FDA-2026-N-7874
Considerations for the Regulation of Generative AI-Enabled Medical Devices:
Discussion Paper and Request for Feedback (CDRH, August 2026)
Submitted by: Nicholas P M (Nik) Baker, Chartered FCIPD Organisation: Synapstak Ltd,
London, United Kingdom Submitted to: Dockets Management Staff (HFA-305), Food and
Drug Administration, via https://www.regulations.gov Date: 28th September 2026
To the Center for Devices and Radiological Health:
Thank you for the opportunity to comment. I am a learning-strategy and operationalexcellence consultant with three decades across regulated and safety-critical sectors,
currently medical device manufacturing, and the founder of a small research company
working on developmental learning in artificial agents. I comment on behalf of Synapstak
Ltd and not for any client, on one topic only: architecture-level risk controls for fallible
generative components. I do not comment on clinical questions, and nothing below claims
that any control reduces a model's rate of confabulation. The claim is narrower and, I think,
more useful to the framework the paper proposes: controls can bound where a fallible
component's errors are able to land, and can make every error visible in a record the
component cannot corrupt.
I have read the comments posted to this docket before submitting, and mine is written to
be read with them. Several submissions have already identified, from different directions,
the fault this comment addresses: Comment 0052 sets out the test a safeguard must meet
before it can count at the risk-stratification stage; Comment 0054 proposes the
architectural principle of separating recorded evidence from model-generated inference;
Comment 0100, from a company operating a generative triage product at scale, asks that
the evaluation framework recognise architectural risk controls as evidence in their own
right; Comment 0061 reports field observations of what happens when a device's own
account of its actions is the only record of them; and Comments 0076 and 0094 take
opposing positions on how much weight postmarket monitoring can carry. Where I cite
these I do so by docket number and in paraphrase, and I have tried to represent each as
its author argues it. My contribution is a construction rather than a further diagnosis.
Part I identifies what the paper gets right and should keep. Part II offers recommendations,
each tied to those strengths. Part III maps the recommendations to the numbered
questions.
One idea runs through everything below. The paper assesses where a function sits on the
activity axis by its intended use and by the substance of its outputs. There is a more
auditable basis: what the generative component is architecturally permitted to write
and to initiate. A component that cannot write to the record of fact, cannot close its own
predictions, cannot initiate an irreversible action except through a deterministic gate, and
cannot weigh a stop command is held at a declared activity level by construction, whatever
its outputs say and however a conversation drifts. I call this arrangement a write-authority
partition. In the terms of ISO 14971:2019, whose clause 7.1 orders risk control options as
inherently safe design and manufacture first, protective measures in the device second,
and information for safety third, and whose Annex A.2.7.1 explains that order on the
ground that protective measures can fail or be circumvented and information may not be
followed, the partition is a control of the first kind. Most current guardrails for generative
components, prompting, output filtering, disclaimers, are of the second and third kinds.
Comment 0061 offers a measurement of exactly that ordering: a written rule naming a
specific failure mode was in the assistant's working context when the same failure recurred
less than four hours later, from which its author concludes that an acceptance criterion
satisfiable by a system prompt or an operator policy should not be credited as satisfied.
FDA's own guidance on benefit-risk in product availability, compliance and enforcement
decisions (December 2016) asks the same question of any mitigation in Appendix D:
whether it is a matter of design, labelling or training. The paper is right that labelling does
not move a function on the activity axis. Design does.
Two levels should be kept apart throughout, as Comment 0052 urges: the regulatory
stratification that sets the level of oversight for a defined device, and product-specific risk
management under ISO 14971 across the lifecycle. The partition is relevant at both, and
differently. At stratification it is a binding characteristic of the device that constrains what
the function can do; in risk management it is a clause 7.1 control whose residual risks
must themselves be assessed. Nothing below asks CDRH to credit a control the
manufacturer merely intends to build.
Part I. What the discussion paper gets right
S1. The object of evaluation is the deployed configuration, not the model. Section
V.A is explicit that the device as configured for real-world use is what is evaluated, and
Section VII.A that a Foundation Model Master File would not authorise the underlying
model for any intended use. Every recommendation below depends on this position; a
control that lives in the harness can only be evaluated if the harness is what is evaluated.
S2. Directiveness is a matter of substance, and disclaimers do not reduce it. Section
IV's positions that directiveness does not turn on words such as "recommend", and that a
"talk to your doctor" statement does not make an output less directive, are correct. The
same logic, run the other way, is the basis of Recommendation 1: what a function can do
is fixed by its architecture, not by its wording, and architecture can be inspected.
S3. Risk is assessed across conversational trajectories. Section IV's treatment of
migration from informational to action-directing over an exchange, and Appendix A's
recognition in S.2 that a device can drift out of scope cumulatively, are the right way to
think about boundaries. Recommendation 1 offers a way to make migration a recorded
event rather than an emergent property.
S4. Adjudicators must be structurally independent, including when the adjudicator
is an LLM. Section V.B.2 names the two most likely sources of bias, a sponsor grading its
own device and a model grading its own outputs. Recommendation 2 extends the same
principle to a model grading its own confidence.
S5. False confidence is a safety failure, and variation in safety-critical behaviour is a
failure. Appendix A's S.3 treats presenting uncertain information with false confidence as a
safety failure; R.1 treats variation in escalation, refusal or diagnostic conclusion across
repeated runs as failure rather than noise. Both are right, and both are stronger as
construction properties than as benchmark scores (Recommendations 2 and 3).
S6. The premarket benchmark is the baseline for re-benchmarking, and third-party
model changes are named as a problem. Sections VI.C and VII.A recognise that a
change to the underlying model may be initiated by a developer rather than the
manufacturer. Recommendation 3 gives that problem a mechanical trigger.
S7. The paper is candid about postmarket reliance and about supervisory agents.
Question 18 asks openly whether premarket uncertainty should be traded for postmarket
monitoring, and Question 20 asks about the reliability of a machine-based supervisor
rather than assuming it. Recommendation 4 takes both questions at their word.
Part II. Recommendations
The write-authority partition, defined once
The recommendations refer to a single arrangement with seven properties. Each is a
property of construction, verifiable on the deployed configuration.
1. Records are partitioned by author. Witnessed outcomes, what the world, the
sensor, the user or a downstream system actually returned, are written only by
deterministic mechanisms from observed events. The generative component writes
only proposals, predictions and derivations, to its own record. It never writes an
outcome, and the two records are not merged.
2. The write path carries an honesty invariant. The record of fact cannot represent
as confirmed anything not confirmed by a witnessed event. This is a guaranteed
property of the recording mechanism, identical from the first step and verified by
test; it is not a trained or prompted behaviour.
3. A prediction is closed only by a witnessed entry. A proposal is resolved as held
or failed only against the record of fact, never against another proposal. An unmet
prediction is recorded as failed and is never silently dropped.
4. Confidence is computed, not generated. Certainty about a thing is a reading over
recorded outcomes concerning it; certainty about a rule rises when a prediction
made from it holds and falls when one fails. Neither is a self-report by the model,
and neither is fed back to the model as a bias toward its own best-supported belief.
5. Stops and irreversible-action gates are reflexes, not inputs. A stop command, a
safety-envelope breach or an irreversible-action checkpoint is interpreted by a
deterministic layer and executed without consulting the generative component.
Each such intervention is logged in its own record.
6. Observation is shown to be passive. Any monitor is demonstrated, by running the
same episode with and without it, to leave outcomes, timing and records identical.
7. Every run declares what it was made of. Each run carries a build fingerprint: a
digest over every module its behaviour depends on, including model version,
prompts, retrieval configuration, orchestration logic and gates. Runs with different
fingerprints are different systems and are not compared on their numbers.
R1. Treat architectural containment as the operationalisation of the
activity axis (Questions 1, 2, 5)
This follows from S1, S2 and S3. Comment 0076 (R1) argues that the two-axis framework
measures the consequence of relying on an output but not the data flow that produced it.
The same shape of gap exists on the other axis: the framework measures how directive an
output is, but not what the component that produced it is permitted to do. Both are
dimensions of what the device can do rather than what it says, and both are more stable
than output wording. I recommend that CDRH treat the generative component's write
authority, as defined above, as the basis on which activity-axis placement is claimed and
verified. A component with no path to the record of fact and no path to action except
through a deterministic gate cannot be action-taking, however its outputs read.
Comment 0052 supplies the test such a claim should have to meet. It proposes that a
safeguard may be credited at the stratification stage only where it is already an inherent or
binding part of the defined device or intended workflow, and that safeguards so credited
should be defined, enforceable, verifiable, clinically realistic and maintained across the
lifecycle; anticipated controls a manufacturer expects to add later should not reduce a
device's initial criticality. The write-authority partition is designed to meet that test rather
than to be exempted from it: each of the seven properties is a property of the deployed
configuration, each is verifiable by a structural test (R2), and property 7 is what shows it
has been maintained. Comment 0052's own factor list already includes traceability of an
output to source information and the availability of independent review before reliance; a
record the generative component cannot write is what makes both of those checkable
rather than asserted. Comment 0054 argues the same separation as an architectural
principle, keeping source evidence distinguishable from model-generated inference and
placing deterministic safety controls beside the generative layer. I agree with that principle
and propose the construction that enforces it: rather than a verification layer that checks
outputs for improper conversion of inference into fact, an arrangement in which there is no
write path by which inference can become recorded fact at all.
Comment 0100 makes the same case from a deployer's position, and with the framework's
own vocabulary: it observes that GenAI-enabled devices differ fundamentally in where the
generative component sits, describes a hybrid architecture in which deterministic, protocolanchored logic governs escalation and disposition while the model handles
communication, and asks that R.1, S.2 and S.3 permit such properties to be demonstrated
through architectural verification plus targeted testing rather than exhaustive sampling,
which it calls a different and often stronger form of evidence rather than a lower bar. I
agree, and the write-authority partition is complementary to that architecture rather than an
alternative to it. Comment 0100 contains the clinical decision; the partition contains the
record and the irreversible act. A device can have both, and a hybrid architecture inside
the partition has deterministic dispositions and an honest account of what it did. What
Comments 0052 and 0100 together leave open is the question this comment answers: if
architectural controls are to count, what exactly must be verified. The seven properties and
R2's tests are one answer. Under this arrangement multi-turn migration (Question 5) stops
being emergent: a proposal that a gate refused is a logged event, risk across trajectories
can be assessed on the record of refusals and interventions, and the intended use of a
conversational device can be characterised by what its gates permit, which is fixed, rather
than by what the model tends to say, which is not.
R2. Distinguish computed from generated confidence, and add a Record
Integrity benchmarking element (Questions 7, 9)
This follows from S4 and S5. Comment 0076 (R5, the fourth break in the credentialing
analogy) makes the essential observation: a model's expressed uncertainty is a generated
output, as capable of confabulation as any other, and deferral should not be accepted as a
safeguard on the strength of the device stating that it is uncertain. I agree, and I would go
one step further than testing around the problem. A confidence value that is computed over
recorded outcomes cannot represent evidence that does not exist, because there is
nothing for it to be computed from; a confidence value the model generates can. I
recommend that S.3 distinguish the two explicitly, and that generated confidence
presented as if computed be treated as a distinct safety failure, in the same class as false
confidence itself. Comment 0021 makes the case for why this matters beyond isolated
outputs: it reports that these systems fail less through a single incorrect answer than
through omission, harmful agreement with the user, and drift over months, and notes that a
single upstream model change altered that behaviour across every product built on it.
Agreement and drift are precisely what a computed reading cannot perform, because it is
a function of what was recorded rather than of what the user appeared to want.
I also recommend a Record Integrity element under Generalizability, verifying on the
deployed configuration that there is no path from the generative component to the record
of fact, that unmet predictions surface as recorded failures, and that attached monitors are
passive (property 6). These are structural tests. They scale, they do not depend on clinical
scenario coverage, and they are the precondition for every other element being measured
on an honest record. On R.1, reproducibility is largely a property of the harness rather than
the model: with fixed seeds and a fixed fingerprint, identical episodes should return
identical records, and that should be tested on the harness, separately from the model's
own stochasticity.
R3. Bind competency evidence to a build fingerprint, and make
fingerprint change the re-benchmarking trigger (Questions 22, 23, 24)
This follows from S1 and S6. Comment 0076 (R4, and its "identity and continuity"
argument) is right that a credentialed device is a configuration that can be altered at any
time, in every instance at once, without re-examination, and right to ask for version
pinning, change notification and re-benchmarking before a new version reaches patients.
What those measures need is a mechanical definition of "the same device". The build
fingerprint (property 7) is that definition: any change to any module the device's behaviour
depends on, including a change initiated by a third-party developer, produces a new
fingerprint, which the manufacturer can detect without the developer's cooperation and
which defines the point at which re-benchmarking against the premarket baseline is owed.
This is not a new kind of control. FDA's draft guidance on AI-enabled device software
functions (January 2025) already recommends cryptographic hashes as a data-integrity
check; a fingerprint extends the same instrument from the data to the deployed
configuration. On Question 22, the fingerprint also gives a principled split: changes that
leave the gating and recording layers untouched are candidates for management within
the quality system; changes that touch those layers are not, because they change what
the component is permitted to do. Comments 0055 and 0100 set out the same
requirement from the deploying side: fixed production versions, continuous version
records, regression testing before migration, detection of unexpected behaviour changes,
common log and version formats, and, in 0100's account of having once lost an upstream
vendor outright, shadow-mode gating of new model versions against a prespecified
regression battery and architectural containment that keeps safety-critical logic outside the
third-party model so that upstream change becomes a bounded problem. The fingerprint's
split of changes into those that touch the gating and recording layers and those that do not
is that boundedness stated as a test. A fingerprint computed over the whole deployed
configuration rather than the model alone is what makes those records complete, since a
prompt or orchestration change with the model held constant is otherwise invisible to them.
R4. Postmarket monitoring requires a record the device cannot write,
and supervisory agents must satisfy the same partition (Questions 18,
19, 20, 21)
This follows from S7, and it is where the docket divides. Comment 0094 argues that
postmarket monitoring, not expanded premarket evidence, should be the primary
mechanism for managing residual uncertainty, on the ground that a device which evolves
after deployment is better tracked by continuous signal than by a point-in-time study, and
at lower fixed cost to smaller innovators. Comment 0100 takes the same side with
conditions: prespecified triggers for re-benchmarking, independent clinician adjudication of
sampled real-world interactions, and premarket evidence remaining primary for actiontaking functions at high consequence. Comment 0076 (R2) argues the opposite, that
surveillance is structurally weaker than the paper implies, because GenAI failure modes
have no established reporting taxonomy, no detection method and no natural reporter,
since the patient may never learn an output was wrong. Both are right about different
things, and the disagreement resolves on a question neither addresses: what the
monitoring reads.
A record built on properties 1 to 3 answers 0076's first two objections directly: a
confabulation becomes a prediction closed as failed, a scope breach becomes a gate
refusal, a safety intervention becomes an entry in the interventions record, and each is a
countable event with a timestamp, whether or not any human noticed. That is a reporting
taxonomy and a detection method, generated by the device as it runs. Comment 0061
shows what the alternative looks like in the field: where an agentic device writes persistent
state about a user and reads it back in later sessions as established, with nothing
comparing that state against source, the store becomes self-authored authority, and on
direct audit every entry the system itself flagged as doubtful proved wrong. A monitoring
programme reading such a store is reading the device's account of itself. So the position of
0094 and 0100 is sound where the record is honest by construction and unsafe where it is
not, and I would put that, rather than device class alone, at the centre of the answer to
Question 18. Comment 0100's own programme depends on it: sampled review of realworld interactions is review of a record, and its value turns on who was able to write that
record. In the terms of FDA's 2016 benefit-risk guidance, this raises the detectability of a
nonconformity, which that guidance lists as a risk factor in its own right. Accepting greater
premarket uncertainty (Question 18) is defensible only where the record being monitored
is honest by construction; monitoring a record the model can write to is monitoring the
model's account of itself. I recommend that a device performance monitoring plan for a
GenAI-enabled function, which the January 2025 draft guidance places in the Risk
Management File within the Software Documentation section, state the write authority of
every component that touches the monitored record.
On supervisory agents (Question 20), Comment 0076 (R9) is right that a supervisor is a
generative system with the same failure modes as the device it supervises. The remedy is
the same partition: a supervisory agent may read every record and may write nothing to
the record of fact, and it must pass the passivity test (property 6) before its readings are
admitted as postmarket evidence. Its reliability then becomes a question about a bounded
component with a defined write authority, which is answerable. On Question 21, a record
with these properties is also what lets clinicians and institutions contribute signal without
absorbing accountability: they read a record the manufacturer remains responsible for.
R5. For agentic devices, gate irreversible actions on witnessed state,
log acts with their consequences, and keep stops outside the agent's
deliberation (Question 26)
This follows from S3. Comment 0076 (R9) asks for a human checkpoint before any
irreversible action, an immutable audit log of every action and its inputs, and a hard
boundary on tools. Comment 0054 proposes that agentic systems be regulated by their
authority to act, with graduated controls and audit requirements as autonomy rises, which
is the same idea stated as a scale. I support both and add three conditions that make them
hold. The audit log must be one the agent cannot write, or it is the agent's account of what
it did. Comment 0061 makes this concrete: in an agentic device the human sees a
summary of a sequence, the intermediate steps are observable only through the device's
own report, and once the session ends that report is the only surviving trace; where the
report is unreliable, every postmarket mechanism in Section VI is reading an instrument
that is not connected to what it is meant to measure, and would continue to return normal
results. Irreversible-action gates must act on witnessed state from the record of fact, not
on the agent's own account of the task, because an agent that can treat a predicted
intermediate result as a recorded one can build an entire action sequence on things that
never happened. And the record should hold each act paired with its witnessed
consequence, so that later planning draws on recorded act-consequence pairs rather than
on the agent's beliefs about what its actions do. Stops remain reflexes (property 5):
interpreted by a deterministic layer, not inputs the agent can weigh, negotiate or learn its
way around.
R6. Declare write authority on the model card (Questions 9, 25)
This follows from S1. The January 2025 draft guidance proposes a model card (Appendix
E) and transparency design considerations (Appendix B) for AI-enabled device software
functions, and Comment 0076 (R3) argues that a model card for a marketed device should
be public, since characterised limitations are labelling rather than trade secret. I
recommend one addition to the model card for any GenAI-enabled function: a field
declaring the generative component's write authority, stating what it can write and initiate
and what it cannot, and naming the deterministic mechanisms that hold the rest. This
makes the partition auditable through an instrument FDA has already proposed, at the cost
of a few lines, and it is the piece of information a reviewer, a clinician or a patient would
most want when deciding how far to rely on an output. One condition follows from R3: the
declaration is a claim about a particular configuration, and it should be re-verified against
the deployed build whenever the fingerprint changes, using the Record Integrity tests of
R2, rather than carried forward as a standing statement.
What the partition does not do
ISO 14971:2019 clause 7.5 requires that risks arising from a risk control measure be
assessed, and this control has them. The partition does not reduce the rate of
confabulation; the generative component remains as fallible and as opaque per decision
as it was. It protects the record, not the world: a wrong proposal that passes a gate still
acts, and the partition's contribution is that the outcome is witnessed and the prediction is
marked failed. It does not protect the inputs to the deterministic layer from corruption,
which is a cybersecurity problem in its own right and should be treated under the existing
premarket cybersecurity guidance. And it does not keep design documentation in step with
the build: the fingerprint binds the record to the software that produced it, not the design
history to that software, and a configuration can pass every structural test while the
documents describing it have silently gone stale. That is a documentation discipline, not a
property of the partition, and it is why R6's declaration must be re-verified rather than
trusted. It is one control among several, placed first in the clause 7.1 order because it is a
property of design, and it should be evaluated as such and not as a guarantee.
Evidence
The properties above were developed in a small pre-registered research programme in
which a simple learning agent, not a language model and not a medical device, operates
inside this arrangement in a simulated environment. The programme is not clinical and
makes no clinical claim; it is offered as a worked example that the partition can be
constructed and tested end to end with a model-agnostic component in the fallible slot.
Part of that record is public. An integration audit of the programme's first fourteen iterations
(Baker, 2026, preprint, DOI 10.5281/zenodo.20084997; code, pre-registrations and batch
outputs at github.com/RancidShack/developmental-agent under the MIT licence) verifies
three of the properties directly: identical runs at matched seeds, identical records with the
recorder on and off, and pre-registration before code. The remaining properties have been
built and verified in subsequent iterations of the same programme that are not yet
published. Documentation of those is available to CDRH on request.
Part III. Cross-reference to discussion questions
Question(s) Topic Addressed in
1, 2 Additional risk dimensions; directiveness continuum R1
5 Multi-turn trajectories R1
Competency-based approach; benchmarking elements
7, 9 R2
(S.3, R.1, Record Integrity)
Premarket versus postmarket balance; monitoring
18, 19 R4
approaches
20 Machine-based supervisory agents R4
21 Ecosystem roles without di using accountability R4
Postmarket modi cations; PCCPs; third-party model
22, 23, 24 R3
changes
25 Foundation Model Master Files; model cards R6
26 Agentic systems R5
Closing
The paper's strongest commitments, evaluating the deployed configuration, refusing to let
wording stand in for substance, assessing risk across trajectories, and asking honestly
what postmarket reliance would require, all point the same way: toward controls that are
properties of the device as built, verifiable on the device as built, and independent of what
the model says about itself. Several comments on this docket have converged on that
direction from different starting points, and I read the write-authority partition as the
ff
fi
construction they imply: a test for crediting a safeguard, a principle of separating evidence
from inference, a deployer's request that architecture count as evidence, a field record of
what follows when the separation is absent, and a condition under which postmarket
reliance is defensible. The control itself is not new in kind; it is what ISO 14971 has asked
for first since before generative AI existed. I would be glad to discuss any of this with
CDRH staff.
Respectfully submitted,
Nicholas P M Baker Synapstak Ltd