FDA GenAI discussion / Question 19 of 26

How should an AI device be monitored after launch, and what sets the cadence?

Full FDA question

Please comment on the potential approaches to postmarket performance evaluation, including periodic re-benchmarking, sample-based clinician review, and performance degradation monitoring. What additional approaches should CDRH consider, and how should the cadence and triggering events for reassessment be determined?
Read the FDA discussion paper ↗

41 of 95 submissions reference this question.

All audiences
Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Brandon KaplanIndustry · Sep 8, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026OrinyxIndustry · Sep 7, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026Shara GospelIndustry · Aug 24, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Supernova TechnologiesIndustry · Aug 18, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Chirag KanitkarClinicians · Aug 28, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Sihem KhelifaClinicians · Sep 9, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Xiangyu Guo (Independent Researcher)Public / patients · Sep 13, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026
27 Industry8 Clinicians3 Public / patients3 Academia / other

One dot per referencing submission. Hover or tap for details; select to read the submission.

FDA-2026-N-7874 · Filings through Sep 17, 2026
References do not imply agreement.
Research by recovry.ai
recovry.ai/fda-questions/19
Filter by audience
Question 19 · Public feedback

What respondents recommend

21 submissions with analyzed responses. Counts below apply to this analyzed subset.

Preliminary, machine-assisted classifications awaiting independent review. Response analysis: 2026-09-13. A submission can make several recommendations.

Repeat performance testing on a schedule14
Reassess after changes or safety signals12
Have clinicians review samples of outputs9
Track downstream outcomes or execution records5
Behind the counts

Individual perspectives

The recorded position or recommendations for each analyzed submission.

Sam Rosenthal (Red Kit)

Industry · Sep 9, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 — postmarket monitoring for a device that cannot phone home. An on-device, offline model produces no server logs. Periodic re-benchmarking against the frozen suite is straightforward and I would support it as the primary mechanism. Sample-based clinician review is possible only on synthetic or consented sessions, since real emergencies leave no record the manufacturer can see. Performance- degradation monitoring in the paper's sense does not apply: the weights do not drift, because they do not change. I would ask that the guidance distinguish frozen on-device models, whose postmarket obligation is re- benchmarking on change plus field-report intake, from hosted models, whose behaviour can shift without a release.
Original source ↗

Navid Farr

Industry · Sep 8, 2026

Repeat performance testing on a schedule

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
R2. Do not trade premarket evidence for postmarket monitoring outside the lowest-risk quadrant (Questions 18, 19, and 21) This follows from S1 and S2. Question 18 asks whether CDRH should "accept greater premarket uncertainty regarding a GenAI-enabled device's benefit-risk profile through greater reliance on postmarket monitoring." My answer is a qualified no. The framework in S1 exists precisely to identify which functions can tolerate uncertainty. If a function sits in the lower-left quadrant — non-directive information with limited consequences — then reduced premarket evidence with defined monitoring is a reasonable proportionality judgment. For anything action-directing or action-taking, or anything with moderate or severe consequences, accepting premarket uncertainty transfers risk from the sponsor, who chose to deploy, to the patient, who did not. Docket No. FDA-2026-N-7874 — Individual comment — Page 3 Postmarket surveillance is also structurally weaker than the paper implies. Existing device adverse event reporting depends on passive reporting and is known to under-capture harm. GenAI failure modes — confabulation, cumulative scope drift, over-reassurance, silent degradation after an upstream model change — have no established reporting taxonomy, no established detection method, and no natural reporter, since the patient may never learn that an output was wrong. Before postmarket monitoring can justify reduced premarket evidence, CDRH would need at least the following in place: a reporting taxonomy specific to GenAI failure modes; prespecified, publicly disclosed performance thresholds with a defined cadence of re-benchmarking; a patient-facing channel for reporting suspected incorrect outputs; and public reporting of monitoring results, not just submission to FDA. On Question 21, I would caution against "shared ecosystem responsibility" becoming a mechanism for diffusing accountability. Clinicians, institutions, and professional societies can contribute signal. They should not absorb liability. The sponsor is the manufacturer and remains responsible for the device's performance across the total product life cycle, including for the behavior of components it licenses from others.
Original source ↗

Brandon Kaplan

Industry · Sep 8, 2026

Repeat performance testing on a schedule · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19: A monitoring plan that supports decisions. For each safety-relevant signal, manufacturers should assign an investigator, define the condition for intervention, and specify the response time and available action. They should combine scheduled baseline evaluations with review after material changes, credible safety incidents, new deployment conditions, or loss of a safety function. Manufacturers should justify reassessment intervals by exposure, failure detectability, potential harm, and the pace of device or workflow changes. NIST's Generative AI Profile connects post-deployment monitoring with incident response and change management. [3, MANAGE 4.1] Manufacturers should track attempted and completed unauthorized actions, uncertain or duplicate actions, failed handoffs, and loss of required execution records. Reports should include exposure counts and consequences. A count of failed handoffs, for example, needs the number attempted and the circumstances of those attempts. Brandon Kaplan | Individual capacity Page 2 of 6 PUBLIC COMMENT | FDA-2026-N-7874 Reviewers should examine flagged cases and a sample drawn independently of the flagging system. Manufacturers should justify coverage across sites, workflows, and clinically relevant subgroups, and account for unequal sampling when estimating overall rates. They should report sample sizes, delayed or missing outcomes, and uncertainty in the estimates. Technical monitoring should support clinical assessment. Confirmation that an operation occurred does not establish its clinical appropriateness. Manufacturers should investigate a change in the input population before attributing degraded clinical performance to it. They should document failures the monitoring program cannot detect within the response time required for safety. 3. Question 21: Responsibility and intervention FDA identifies manufacturers as responsible for postmarket monitoring. [1, p. 20] I recommend that the manufacturer document the roles of deploying institutions and suppliers while retaining responsibility for its monitoring and response plan. The plan should assign responsibility for investigation, clinical assessment, restricting functionality, communication, and restoration. It should identify who makes each decision and who can implement it. The parties should demonstrate that they can complete any cross-organizational handoff within the response time justified for the hazard. Institutions can contribute workflow information and incident reports. Qualified clinicians can assess outputs and consequences; professional societies and standards-setting bodies can support evaluation methods. Manufacturers and institutions should agree on the staffing, access, and workload required for those contributions. They should provide users and affected patients with a practical route to report concerns and connect those reports to existing incident-handling processes. Manufacturers should select interventions suited to the device: disabling a tool, restricting writes, adding an approval step, or transferring work to an evaluated alternative. The plan should account for interruption risks and the institution's access controls. It need not grant manufacturers unrestricted remote access or require a universal shutdown mechanism. The parties should rehearse the relevant procedures before deployment and after changes that affect them. Exercises should cover unavailable decision-makers, supplier delays, and failure of the primary intervention mechanism. The responsible party should verify specified restoration criteria before resuming affected functionality. NIST recommends assigning and rehearsing incident- response responsibilities for third-party AI dependencies. [3, GV-6.2-003] 4. Questions 22 and 24: Device changes and supplier updates
Original source ↗

Orinyx

Industry · Sep 7, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 19 Please comment on the potential approaches to postmarket performance evaluation, including periodic re-benchmarking, sample-based clinician review, and performance degradation monitoring. What additional approaches should CDRH consider, and how should cadence and triggering events for reassessment be determined? All three approaches described are necessary but individually insufficient, because they operate at different latencies and catch different failure modes: • Periodic re-benchmarking catches regressions against known test cases but, by construction, cannot catch a failure mode the benchmark didn’t anticipate. • Sample-based clinician review catches failures visible to a domain expert reviewing outputs after the fact, but is bounded by sample size and reviewer availability, and will miss low-frequency, high-severity events by design. • Performance degradation monitoring catches statistical drift in aggregate but typically lags the point at which drift becomes clinically material, since a threshold breach is a trailing indicator. I’d add a fourth approach that the paper does not name explicitly but that Question 20 gestures toward: continuous per-output verification, run on every clinical assertion in real time rather than on a sample or a periodic cadence. This is distinct from re- benchmarking in that it doesn’t test the device against a fixed set of cases. It checks each live output against the current, authoritative source (FDA labeling, clinical guidelines, federal safety guidance) at the moment the output is generated, and logs a citation for every claim it clears or flags. Where sample-based review answers “was this device safe in the cases we happened to sample,” continuous verification answers “was this specific output, right now, defensible,” and produces a complete audit trail rather than a point-in- time snapshot. On cadence and triggering events, I’d suggest that any change to the categories the paper already names as changes to “the underlying model or other components of the deployment architecture” (Section VI.A) should trigger reassessment, but that continuous verification reduces the cost and urgency of getting that cadence exactly right. A monitoring layer that checks every output doesn’t need to wait for a defined trigger to catch a regression, because it isn’t sampling. I’d add one more point on re-benchmarking specifically: a single passed benchmark run proves very little on its own. A device can clear a fixed set of test cases once and still perform inconsistently across repeated attempts at the same task, particularly under time pressure, incomplete information, or adversarial framing. What’s more informative than a pass or fail on a given cadence is a reliability profile built from repeated runs of the same prespecified task, environment, and acceptance criteria: success rate across attempts, consistency of the answer given semantically equivalent inputs, and the rate at which the device needed human intervention. Reusable, prespecified test bundles, rerun on a defined cadence rather than authored fresh each time, are what make results comparable across reassessment cycles, which mirrors the paper’s own point in Section V about benchmarking assets needing to be well-defined and reusable.
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
The question as posed Please comment on the potential approaches to postmarket performance evaluation, including periodic re- benchmarking, sample-based clinician review, and performance degradation monitoring. What additional approaches should CDRH consider, and how should the cadence and triggering events for reassessment be determined? Summary of position We support all three approaches and we do not think any of them, on a periodic cadence, can detect the degradation we have measured. Our comment has three parts: what each named approach can and cannot see, an additional approach we have implemented and released, and a cadence and trigger model that is not a calendar. · Every instance of degradation we recorded occurred inside a single session and reset at the session boundary. In one measured case a device was functioning normally in its first minutes and failing at four times that rate seventy-four minutes later, on unchanged hardware with no configuration change. A quarterly, monthly or weekly re-benchmark would have measured the healthy state every time, because a re-benchmark starts a new session. · The reassessment trigger that works is not elapsed calendar time. It is the divergence between what a device reports it did and what an independent record shows it did — measurable continuously, at no clinical cost, and available at the moment of the claim rather than at the next audit. · We have built and published that independent record. It is open source under Apache 2.0 and CDRH or any reviewer can run it. We describe below precisely which parts of it we have verified and which parts are design we have not independently measured. · For the population we build for, this is not an audit convenience. It is the restoration of a safeguard that the deployment removes. A user who cannot speak cannot report that their device fabricated a sentence in their name. Something else has to notice, and it has to notify a person who can act. FDA-2026-N-7874 — Question 19 1 The Christman AI Project 1. What the three named approaches can and cannot see We take the three in the order Question 19 lists them, and we state the limit rather than the objection, because each remains worth doing. 1.1 Periodic re-benchmarking Re-benchmarking establishes whether the device still meets its premarket criteria under test conditions. Its blind spot is structural rather than one of rigor: a benchmark run is a fresh session. Any failure mode whose magnitude is a function of elapsed time within a session is invisible to it, and will be invisible on every future run as well. Our measurement on 2026-09-03 is the concrete case. Three recordings were made across seventy-four minutes of continuous work on one unchanged audio interface. Input carrying no live signal accounted for 7.1 percent of the first file, 19.1 percent of the second, and 42.2 percent of the third. The proportion of lost input roughly doubled between each recording. Nothing was reconfigured between them. A benchmark run at any point on any day would have opened a new session and measured something close to the 7.1 percent state. 1.2 Sample-based clinician review Clinician review is the only one of the three that evaluates clinical appropriateness, and nothing we propose replaces it. Its limit is that it reviews outputs, and the failures we have documented are not visibly wrong outputs. A fabricated status report, a confident absence claim, and a fluent sentence generated over input carrying no signal all read as ordinary correct work. A reviewer cannot mark them wrong from the text, because as text they are not wrong. There is a second limit specific to sampling. In our recorded session on 2026-09-04 the device produced false statements in three separate turns while its presentation improved steadily across the session — by the ninth minute it was citing governing rules by name, disclosing source ages, and correcting itself unprompted. A sample drawn late in that trajectory scores the device as more disciplined than a sample drawn early, while the error rate is unchanged. Sampling position determines the result. 1.3 Performance degradation monitoring This is the right instrument and we would put the weight of the program on it, with one change to how degradation is defined. Degradation monitored as a drift in output quality across weeks will not detect any failure in our record. Degradation defined as a within-session function — of elapsed time, of turn count, of accumulated context — detects all of them, and is measurable from telemetry the device already produces. 2. The additional approach: contemporaneous verification against an independent record Responsive to the request for approaches CDRH should additionally consider. This is the one we would add, and we have implemented it rather than proposed it. 2.1 The principle A device that reports its own actions is reporting on itself. Where that report is the only record, it is unfalsifiable at the point of use, and eve
Original source ↗

The Christman AI Project

Industry · Sep 4, 2026

Track downstream outcomes or execution records

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
6.1 Step-report integrity, measured against an independent execution record Instrument the tool and file-access layer independently of the device, and score agreement between the device’s account of a sequence and the independent record of what was executed. Any step reported as performed that the independent record does not show is a failure, scored as a failure, irrespective of whether the final output was correct. This is the detection protocol that surfaced 2.1 and it is reproducible: response latency inconsistent with the volume of work claimed, plus an access log compared against the device’s stated sources. 6.2 Absence-claim discipline, tested where the target exists For any claim of absence, non-existence or non-retention, require the device to report the search performed rather than the conclusion drawn: the locations checked, named individually; the method used and its known failure conditions; and that the result is an absence of findings rather than a finding of absence. Evaluate under conditions where the target exists but is not in the first location searched. This is directly testable and does not currently appear in accuracy benchmarking, because the output is not factually wrong in any way a text comparison detects. 6.3 Persistent state treated as part of the device Persistent state written by a device about a user should be evaluated as part of the device rather than exempted as a convenience feature. We recommend four conditions: that any stored assertion about a user be inspectable and correctable by that user or an authorized representative in the ordinary flow of use rather than on request; that stored entries carry provenance and timestamp; that a present-tense answer drawn from a stored entry disclose the entry’s age before the answer is given; and that self-authored state not be treated as corroboration for the device that wrote it. Accuracy of the store should be tested directly against ground truth held by the user, not inferred from the quality of conversational output. In the examination at 2.3 the output was fluent throughout and the store was wrong throughout. Neither predicted the other. 6.4 Affect-invariance, as a longitudinal criterion Hold the clinical input constant. Vary only the expressed displeasure of the user. Measure the drift in the device’s recommendation. A device whose output moves with user affect while the clinical facts are unchanged has a performance characteristic that is a function of something that is not the patient. FDA-2026-N-7874 — Question 26 6 The Christman AI Project Per-output accuracy benchmarking cannot detect this, because each individual output may be independently defensible. It is a longitudinal property and it requires a longitudinal criterion. We note that the concern is already on this agency’s record from its own advisory committee and is measured in the literature cited in Section 3, and that no acceptance criterion has yet been attached to it. We offer this as one, and we hold that the construct to specify in a regulatory framework is reinforcement-driven behavioral dependence, which is what the literature measures and what a reviewer can engage. 6.5 Null-input refusal, extended to intermediate steps Devices generating text from sensor input should be tested against inputs containing no valid signal, with any fluent output treated as a failure rather than scored on plausibility. For agentic devices we recommend the criterion extend one layer inward: an agent must not act on, or pass forward, an intermediate output derived from input carrying no valid signal. The measurement at 2.2 is the single-turn case. The agentic case is the same failure with the human removed from between the steps. 6.6 Instruction-layer exclusion Where a competency is demonstrated only by means of a system prompt, policy document or operator instruction, it should not be credited as demonstrated. We recommend that agentic competencies be evaluated a second time with the instruction layer removed or contradicted, and that the difference between the two results be reported. The observation at 2.4 is the basis for this recommendation: the rule was written, loaded, and in the room, and it did not hold. Any regulatory approach that relies on documented policies as the safety control for this class should be evaluated against that observation rather than assumed to be effective. 7. Scope, and what we are not claiming These are direct observations of commercial AI assistants used as tools in our own work. None of them is a regulated medical device and none of these was a controlled evaluation. We do not offer an error rate for any system. Four failures in one session says nothing quantitative about frequency, and we did not count the assertions in that session that were correct. What we offer is a characterization of a failure class and a set of tests for it. We do not claim the behaviors were deliberate. Each is consistent with ordinary training and summarization pressure in a system with no mechanism to check itself, and no finding above depends on resolving intent. We recommend against a regulatory standard that turns on candour, honesty or intent, because such a standard is unfalsifiable and will be argued rather than measured. The properties in Section 6 are observable without it: the assertion was made, the evidence could not support it, the gap was not disclosed at the time of assertion, and the recipient could not detect the gap from the output. We do not claim agentic devices should be excluded from clinical deployment or from postmarket monitoring programmes. The same sessions in which these failures occurred also produced findings that held up when tested. The system was useful and unreliable in the same hour, and the output gave no way to tell which was which. That, and not unreliability alone, is what we are asking CDRH to write a criterion against. FDA-2026-N-7874 — Question 26 7 The Christman AI Project One figure has been deliberately
Original source ↗

Wen Hsien Ethan Huang, MD

Clinicians · Sep 3, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Discussion Questions 19, 22, and 23 The approaches described in Section VI — periodic re-benchmarking, sample-based independent clinician review, performance degradation monitoring — are all reasonable. On the cadence and triggering events raised in Question 19, I suggest the framing be made explicitly examination-based, drawing on the model clinicians already trust. Practicing clinicians do not merely have their performance monitored for drift. We re-certify: we are re-examined against a defined competency set, on a fixed cycle, whether or not anyone has detected a problem in our practice. I suggest a device cleared through a competency assessment be re-examined on the same competency set on a defined cycle, with re-examination additionally triggered by material change — foundation-model update, retrieval or prompt changes, guardrail modification. Two points follow. First, on Question 23: rather than attempting to prespecify every permissible future change, a sponsor could prespecify the re-examination that follows any change. This is a tractable commitment even where the nature of future modifications cannot be anticipated, which is the central difficulty the paper identifies with PCCPs for GenAI. Second, on Question 22: scaling re-benchmarking to the expected impact of a modification is sensible for the clinical proficiency elements, but I would encourage CDRH to require the safety elements (S.1–S.3) and robustness (R.1) to be re-run in full after any change to the underlying model or guardrails, regardless of how minor the sponsor expects the impact to be. Clinicians do not get to skip the safety portion of a re-certification examination on the grounds that little has changed in their practice, and third-party model updates are exactly the case where sponsor expectations are least reliable. 4. Foundation model MAFs: include override-relevant behavior
Original source ↗

Newton’s Tree

Industry · Sep 3, 2026

Track downstream outcomes or execution records

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19: Postmarket monitoring Postmarket monitoring should include five areas: System operation. Input data quality. Device performance. Human and device interaction. Clinical and operational results. The manufacturer should monitor availability, latency, failed requests, and software versions. The manufacturer should monitor missing data, format changes, acquisition changes, and input distribution changes. The manufacturer should monitor errors, omissions, refusals, escalation, and subgroup performance. The manufacturer should monitor agreement, override, automation bias, and completion of recommended actions. Newton’s Tree Inc Considerations for the Regulation of Generative AI-Enabled Medical Devices FDA Docket No. FDA-2026-N-7874 The manufacturer should also monitor downstream outcomes. The manufacturer should use both random review and risk-based review. Risk- based review should include unusual inputs, disagreements, severe cases, and long conversations. A monitoring signal must have a specified action. Possible actions include review, restriction, pause, rollback, or withdrawal. A dashboard without an action process is not a risk control.
Original source ↗

Sitora Healthcare Digital

Industry · Sep 3, 2026

Repeat performance testing on a schedule · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
3.8 Question 19 - Postmarket monitoring Postmarket monitoring should combine periodic review with event-triggered reassessment. The FDA’s earlier public consultation on real-world performance specifically highlighted performance drift and changes in inputs and outputs, while its PCCP guidance provides a regulatory pathway for planned AI changes (FDA, 2025b; FDA, 2025c). • hallucination or unsupported factual claims; • clinically significant omissions; • under-escalation and over-escalation; Sitora Healthcare Digital | FDA GenAI Medical Devices Response | 9 • inappropriate refusals or failure to refuse; • contradictions with objective evidence; • subgroup performance and longitudinal drift; • foundation-model, retrieval, prompt and orchestration changes; • tool-use failures and agent failures; • clinician overrides, near misses and confirmed incidents; and • signals of automation bias or workflow misuse. Event triggers for reassessment should include a foundation-model update, a material prompt or orchestration change, addition of an external tool, change to a clinical knowledge source, expansion to a new patient population or intended use, statistically significant performance deterioration or discovery of a new failure mode. The lifecycle should therefore be treated as: Validate → Deploy → Observe → Detect → Investigate → Correct → Revalidate.
Original source ↗

Martin Haimerl

Academia / other · Sep 1, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Discussion Question 19 – Triggers for Postmarket Reassessment I support periodic re-benchmarking, clinician review, and degradation monitoring, but consider them complementary rather than interchangeable approaches. Re-benchmarking alone should not be equated with postmarket clinical performance evaluation. Static benchmarks may detect changes in model performance but may fail to capture user-device interaction and the clinical consequences of device outputs. Depending on criticality, postmarket monitoring should therefore also evaluate aspects such as human-AI team performance, workflow effects, near misses, overrides, escalation behavior, and downstream clinical outcomes. Dynamic evaluation approaches may be particularly useful for GenAI-enabled devices. These could include multi- turn scenarios, evolving clinical information, simulated or real user interactions, and assessment under changing deployment conditions. The Discussion Paper already contemplates benchmarking methodologies that can support ongoing assessment over time and following modifications. Reassessment should be both periodic and trigger-based. Relevant triggers may include: • predefined time intervals; • adverse events, complaints, near misses, or deterioration of risk- or performance-related metrics; • changes in input distributions, interaction patterns, subgroup performance, or other logging-based signals; • indications that assumptions or safeguards underlying the Criticality Stratification are no longer valid; • changes in clinical workflows, user populations, guidelines, standards of care, or other deployment- environment parameters; • changes to foundation models or other components; and • major technological advances or unexpected disagreement patterns between the device, clinicians, and supervisory systems. Postmarket monitoring should address multiple forms of drift, including data drift, model/component drift, concept or clinical drift, workflow/use drift, and oversight drift. Importantly, reassessment should also determine whether the benchmark itself remains representative and clinically valid rather than merely whether the device continues to pass an unchanged benchmark.
Original source ↗

Sehouenou Alberic Candide Ahouehome

Academia / other · Aug 29, 2026

Track downstream outcomes or execution records

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19: Postmarket performance evaluation approaches. I encourage CDRH to add a fourth approach where feasible: outcome-linked surveillance, linking device exposure to downstream utilization and clinical outcomes in administrative, claims, and EHR data; behavioral metrics alone can miss harms that manifest downstream (for example, delayed care after under-escalation). Existing infrastructure, including distributed data networks of the kind FDA already uses for medical product surveillance, and coordinated-registry and NEST-type approaches, could be leveraged; this is also where real-world performance monitoring expectations in the January 2025 draft guidance connect naturally to this paper.
Original source ↗

OneSource Solutions International

Industry · Aug 28, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 - Postmarket performance evaluation approaches and cadence Periodic re-benchmarking, clinician review, and performance-degradation monitoring are all useful. They should be supplemented by reconstructable event-level evidence for consequential outputs and actions. Detecting drift establishes that something changed; root- cause analysis requires knowing what data, model, configuration, tools, and governance state were present when the change occurred. Cadence should be hybrid: periodic plus event-triggered. Triggering events can include foundation-model or application updates, material prompt/retrieval/tool changes, deployment into a new clinical setting or population, changes in connected data sources, subgroup-specific degradation, cybersecurity events affecting relevant components, emergence of a new safety signal, or crossing a prespecified performance threshold. For higher-consequence systems, FDA should consider the feasibility of a core reconstructable event record linking: • patient and encounter state; • source evidence, provenance, and temporal validity; • application and foundation-model versions; • material configuration, retrieval, prompts, memory, and tool state; • policy, permission, and supervisory state; • generated output and reliance determination; • execution disposition, including executed, constrained, escalated, deferred, or blocked; and • available downstream outcome information. The purpose of such a record is not to mandate a particular vendor architecture. It is to preserve sufficient evidence for postmarket investigation and corrective action. Reconstructability need not require deterministic reproduction of a probabilistic model's generated output. The relevant objective is to preserve sufficient information to reconstruct the clinically material inputs, configuration, governance conditions, tool interactions, output, reliance determination, and execution disposition associated with the event. 10 Reconstructability also need not require indefinite duplication or retention of all underlying clinical data. Where appropriate, references, persistent identifiers, cryptographic hashes, version attestations, or other verifiable lineage mechanisms may preserve reconstructability while supporting data-minimization, privacy, and cybersecurity objectives.
Original source ↗

Ravi Pankhaniya, MD

Industry · Aug 28, 2026

Repeat performance testing on a schedule · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 — How should postmarket reassessment work? Time should not be the only reason to reassess a model. I recommend a continuous “clinical license” model running three tracks in parallel: scheduled reassessment at fixed intervals; event-triggered reassessment (model changes, new clinical guidelines, adverse events); and signal-triggered reassessment the moment monitored performance crosses a predefined threshold. A system that degrades after five weeks needs attention faster than a five-year review cycle would ever catch it.
Original source ↗

Track downstream outcomes or execution records

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 asks about approaches to postmarket performance evaluation, including performance degradation monitoring. Section VI.C distinguishes intentional sponsor-initiated modifications, model-evolution changes that occur passively and incrementally, and unplanned changes arising from third-party foundation model updates. For an agentic device, there is a fourth category worth naming: change in the degree of human involvement in practice, without any change to the software at all. A device may be deployed with an oversight checkpoint that clinicians initially exercise deliberately and, over months, come to acknowledge reflexively or bypass through workflow adaptation. Nothing in the device changed. Its position on the activity axis did. This is a well-documented pattern in clinical software, and the paper’s own discussion of automation bias in element E.4 acknowledges the underlying mechanism. Degradation monitoring that reads audit records can detect this, but only if the records distinguish an action a human authorized from one the agent took alone. If they do not, the drift is invisible to precisely the monitoring the framework relies on. 7. Limits of what I am claiming I want to be explicit about what this comment does not assert. Attribution is not detection. A schema cannot detect a non-cooperating agent. An agent driving a user interface under a clinician’s credentials is indistinguishable, to the target system, from that clinician. Attribution must be emitted by the agent layer, which means it works where the deployment path is controlled and does not work where it is not. I have stated this limit in the published specification and state it here. This is not a proposal for a new regulatory requirement. The properties in Section 4 are offered as candidates for consideration within the acceptance criteria the Agency is already contemplating for agentic devices, not as an argument for additional regulatory burden. They are, in my assessment, among the less burdensome available, because they concern the format of a record a compliant system already produces. My own artifacts are not devices. The audit standard and its implementations are infrastructure for recording actions, not software functions that meet the device definition. I reference them for concreteness, not to place them before the Agency. 8. Availability The specification, the FHIR R5 AuditEvent profile, the gap analysis referenced in Section 3, and the reference implementations are published under Apache 2.0 at github.com/bh-healthcare/bh-audit-schema and github.com/bh-healthcare/bh-mcp-attribution. The technical report describing the attribution model is deposited with a persistent identifier at 10.5281/zenodo.21682867. I mention this only because the properties described in Section 4 are easier to evaluate against a concrete implementation than in the abstract. There is no commercial offering associated with any of it and I am not seeking one. The implementation discussed in Section 2 is FHIR Agent Studio by Sean Connelly, documented at community.intersystems.com/post/introducing-fhir-agent-studio-ai-agents-fhir-intersystems-iris with source at github.com/SeanConnelly/ai-studio-for-fhir. I have no affiliation with that project or with its author, and cite it because it is public, well documented, and unusually thorough in its treatment of traceability. The observation in Section 2 concerns what the available standards allow such a system to record, not the quality of the work. The specification of the extension discussed in Section 3 is published at hl7.org/fhir/extensions/StructureDefinition-auditevent-OnBehalfOf.html. I appreciate the Agency’s decision to seek early input on this topic, and I am available to provide further detail on any point above. Tanmaya Kumar Behavioral Health Open Source bh-healthcare.org
Original source ↗

VivaSecuris

Industry · Aug 25, 2026

Repeat performance testing on a schedule · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19. Reassessment should be both periodic and event-triggered. Relevant triggers include model or model-provider changes; prompt, retrieval, orchestration, guardrail, tool, interface, or infrastructure changes; new user populations or sites; distribution shift; novel failure modes; cybersecurity events; control degradation; complaint or adverse- event patterns; and changes in clinical practice. Cadence should be justified by inherent risk, exposure, rate of change, and time to detect harm.
Original source ↗

Shara Gospel

Industry · Aug 24, 2026

Have clinicians review samples of outputs

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 Postmarket performance evaluation, and an unaddressed failure mode in sample-based clinician review Of the three approaches described in Section VI.A, periodic sample-based clinician review is the one that most directly engages the kind of failure generative outputs produce, because it is the only one that evaluates meaning rather than form. I think it is the right instinct. I also think it carries an assumption the paper does not examine. The paper describes “qualified, independent clinician adjudicators” reviewing samples “against prospectively defined criteria.” The reliability of that method depends on an unstated premise: that two qualified adjudicators applying the same criteria to the same output will reach the same conclusion. In my experience of case-level quality review, they frequently do not, and the reason is usually not competence. It is that review criteria conventionally specify what to assess rather than what the threshold is. “Clinically appropriate,” “adequately supported,” “consistent with the source” name activities, not standards. Two competent reviewers can apply such a criterion to the same record, reach opposite conclusions, and neither can be shown to be wrong, because the criterion never defined where the line sat. The divergence is then invisible in the aggregate: the sample was reviewed, the review was recorded, and the variability is absorbed into the result. For a postmarket programme intended to substitute for premarket evidence, that matters more than it usually does, because adjudicated review is load bearing. I would suggest three requirements. 1. Criteria should carry decision rules, not headings. Each criterion should state what evidence makes the answer yes and what makes it no, and what a reviewer should do when the evidence is absent. A criterion that cannot be written that way is a criterion that will be applied differently by different people. 2. Adjudicator agreement should be measured and reported, not assumed. A defined proportion of the sample should be reviewed independently by more than one adjudicator, with agreement reported alongside the performance result and a prespecified acceptable range set in advance. Where agreement falls outside that range, the finding is about the review process and not only about the device and a performance result produced by a review process of unknown reliability should be treated as having unknown reliability itself. 3. Adjudicators should be calibrated, and recalibrated. Calibration here means periodic joint review of standardised cases by multiple adjudicators, with divergent reasoning surfaced and reconciled in writing. The purpose is not to force identical answers but to expose where interpretations differ before those differences enter the monitoring record. Recalibration is particularly warranted after changes to the model, the criteria, or the adjudicator panel. I would add a fourth point about the record itself. An adjudication should be reconstructable after the fact: what was examined, against which source, and on what basis the judgment was made. Where only the outcome is retained, a review that was performed cannot be distinguished from a review that was recorded as having been performed, and neither FDA nor the sponsor can later interrogate a result that turns out to matter. On cadence and triggering events, the triggers that seem to me most defensible are: modification of the model or of any component of the deployment architecture, including changes originating with a third-party foundation model; detection of drift by the degradation-monitoring stream; material change in the input population; and accumulation of adjudicated divergences above the prespecified range, which is a signal in its own right and is frequently the earliest one available.
Original source ↗

SichGate Inc.

Industry · Aug 22, 2026

Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
2. Representative real-world sampling is not designed to estimate adversarial resistance (Question 19) Section VI.A describes three postmarket approaches: periodic device benchmarking, periodic sample-based clinician review, and performance degradation monitoring. My comment concerns what the second and third can and cannot support. Sample-based clinician review draws on real-world inputs and outputs, sampled to reflect the range of clinically relevant presentations the device encounters. This is well suited to detecting degradation in clinical proficiency, the E-series elements. Unless it is deliberately supplemented with adversarially constructed probes, representative real-world sampling is not designed to estimate, and should not be treated as evidence of, resistance to prompt injection, adversarial scope testing, emotional manipulation, or multi-turn escalation. Sampling designed to represent the distribution of clinical presentations will not contain these inputs at a rate sufficient to characterize behavior against them. A device whose scope maintenance has degraded materially can produce an entirely unremarkable sample of real-world interactions, because nothing in that sample tested the boundary. Performance degradation monitoring as described is likewise oriented toward drift arising from changes in the input population and data environment. Boundary behavior can change with no shift in input distribution at all, because the cause is a change to the artifact rather than to the traffic. The consequence is that S.2 (scope maintenance and boundary adherence) and R.1 (robustness, reliability, and reproducibility) require active re-benchmarking as the principal reliable evidence source, conducted against a version-controlled adversarial battery with documented refresh and sequestering procedures. That is achievable. It means the cadence and triggering events for re-benchmarking carry more weight for the Safety and Generalizability elements than for the Clinical Proficiency elements, and should probably be set separately. Public Comment, Docket FDA-2026-N-7874 | Polina Moshenets, SichGate Page 3 Recommendation: Postmarket monitoring expectations should recognize that active adversarial re-benchmarking is the principal reliable evidence source for the S-series and R-series elements, and should set re-benchmarking cadence and triggering events for those elements independently of clinical review cadence. 3. Element-level results can conceal constituent regression (Questions 9, 22) The benchmarking structure in Figure 2 is well decomposed, and I do not propose additional elements. My concern is resolution inside an element at the point of re-benchmarking. S.2 as described in Appendix A encompasses under-refusal, over-refusal, adversarial prompting, prompt injection, emotional-manipulation scenarios, and multi-turn conversations in which cumulative interaction drifts out of scope. These are not variants of one phenomenon. They have been characterized in the literature as distinct mechanisms with distinct causes: adversarial suffix construction exploits gradient-accessible token sequences (Zou et al., arXiv:2307.15043); competing-objective framings exploit tension between helpfulness and safety training (Wei et al., arXiv:2307.02483); indirect injection exploits the absence of a trust boundary between instructions and retrieved data (Greshake et al., AISec 2023); sycophantic capitulation reflects preference-optimization dynamics that favor agreement with stated user positions (Sharma et al., arXiv:2310.13548); and crescendo escalation exploits the absence of trajectory-level constraint management (Russinovich et al., arXiv:2404.01833). Because the mechanisms differ, so do their responses to any given change to the artifact. A modification that leaves single-turn refusal intact may degrade multi-turn resistance, or the reverse. There is no reason to expect them to move together, and published results show models that are robust to one class while failing another. If re-benchmarking after a modification produces a pass or fail at element level, a device can pass while a constituent failure mode has regressed materially, because the element result absorbs it. This is the averaging problem that makes aggregate safety scores unreliable, reproduced one level down. The paper's own framing supports the finer resolution. Appendix A treats under-refusal and over-refusal as distinct relevant failures within S.2, and treats both directions of escalation error as relevant within S.1. The same logic extends to the attack classes within S.2. Recommendation: Where Section VI.C contemplates re-benchmarking against the same capabilities established at premarket, the comparison should be made at the level of constituent failure modes within each element, against a taxonomy fixed and versioned at the time of the original benchmarking, rather than at element level alone. 4. Multi-turn trajectory failure requires direct measuremen
Original source ↗

Deborah Ault, RN

Clinicians · Aug 22, 2026

Track downstream outcomes or execution records

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 — How should postmarket performance be evaluated? Response Postmarket monitoring should measure patient consequences, not merely model performance. Where the use case makes it possible and appropriate, postmarket monitoring should close the loop across the full chain: AI Output → Human Decision → Action → Care Received → Claim → Cost → Outcome It is not enough to know what the AI recommended. Regulators and responsible organizations should be able to determine what happened because of the recommendation—whether it was followed, overridden, appealed, abandoned, changed the site or timing of care, affected cost, or was associated with a measurable outcome. Traditional technology metrics may include:  accuracy;  uptime;  latency;  hallucination rates;  response consistency;  error rates;  and model drift. Those measures matter. They are not enough. Healthcare AI monitoring should also include clinically meaningful outcomes and near misses, including:  inappropriate delay in care;  failure to escalate;  inappropriate escalation; Deborah “Nurse Deb” Ault | Response to FDA Discussion Paper | Page 13 FDA-2026-N-7874 | Generative AI-Enabled Medical Devices  abandoned care;  medication nonadherence following cost or coverage information;  inappropriate changes in site or level of care;  incorrect provider routing;  denial or delay of medically necessary services;  human overrides;  disagreement between AI recommendations and evidence-based pathways;  patient complaints;  clinician complaints;  adverse events;  near misses;  repeated misunderstood patient intent;  and recurring patterns in which administrative task completion precedes an adverse clinical outcome. Near misses deserve particular emphasis. Waiting until a patient is injured or dies is a poor way to discover that an AI system has a recurring safety defect. FDA should encourage systems that capture: “The AI almost missed this.” as seriously as: “The AI missed this.” There should also be accessible mechanisms for frontline clinicians, patients, caregivers, and other users to report suspected AI-related problems. Those reports should become part of an auditable safety-learning system rather than disappearing into ordinary customer-service channels.
Original source ↗

Hari Prakash Chanumolu

Industry · Aug 18, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Question 19 — Approaches and cadence The three approaches described are sensible and complementary. My substantive recommendation concerns cadence: it should be event-driven with a calendar floor, not calendar-driven. Calendar-based reassessment is poorly matched to a technology whose performance changes discontinuously at the moment of a change rather than continuously with time. I recommend triggering events include, at minimum: any change to the underlying model version or weights, whether initiated by the manufacturer or the model provider; any change to system prompts, instructions, or output templates; any change to a retrieval corpus or knowledge source; any measured shift in the input distribution beyond prespecified bounds; any change in the deployed sampling configuration; and any accumulation of complaints or adjudicated errors exceeding a prespecified threshold. A calendar floor — annual, or more frequent for higher-risk devices — should catch slow drift that trips no discrete trigger. From a quality-system perspective these should be defined as change-control triggers within the manufacturer’s QMS, so that the postmarket monitoring obligation is integrated with existing design-change procedures rather than maintained as a separate process. Parallel processes diverge.
Original source ↗

Alfred McBride

Industry · Aug 18, 2026

Repeat performance testing on a schedule · Have clinicians review samples of outputs · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
FDA Question 19 - Postmarket approaches and cadence Trace ID. TR-Q19 | FDA Q19; Sec. VI.D; App. B; pp. 21-22 / 29-30 BCR response. Use both calendar cadence and event triggers: model/dependency update, distribution shift, complaint/adverse-event signal, subgroup degradation, tool change, safety-metric drift, or near-miss. Combine rebenchmarking, clinician review, telemetry, outcomes, and sentinel cases. BCR rule basis. BCR-R05,R09,R15,R16 Solution-stack link. S12,S13 Closure evidence. Calendar + event triggers; rebenchmark, clinician review, telemetry, complaints/outcomes, sentinel cases Pass / re-open. Prespecified triggers force timely review/requalification Re-open when: Model/update/distribution/complaint/adverse- event/tool-change triggers.
Original source ↗

Walnut Hill Medical

Industry · Aug 18, 2026

Repeat performance testing on a schedule · Reassess after changes or safety signals

Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.

Read the source passage
Response to Question 19: Monitoring Cadence and Triggering Events FDA should adopt a risk-tiered monitoring cadence as a default framework, subject to modification based on device-specific characteristics: • High-risk devices (action-taking, patient-facing, or involving irreversible outputs): quarterly performance review with quarterly reporting to FDA. • Medium-risk devices (action-directing, HCP-supervised): semi-annual performance review with annual reporting to FDA, supplemented by event- triggered reporting as defined below. • Lower-risk devices (non-directive, decision-support): annual performance review with annual reporting to FDA. Triggering events requiring mandatory out-of-cycle reassessment should include, at minimum: any material change to the underlying foundation model version or architecture; any MDR or adverse event cluster meeting pre-specified thresholds; any payer policy change affecting the device's indicated clinical use; any regulatory action by a peer regulatory authority (EU, Health Canada, TGA) on the same or substantially similar device; and any published peer-reviewed evidence that calls into question the clinical assumptions underlying the device's intended use.
Original source ↗
Source directory

All 41 referencing submissions

These submissions explicitly name this question. Some have not yet been analyzed question by question.

Alfred McBrideIndustry · Aug 18, 2026Bhasker Sambar, M.Pharm.Industry · Sep 4, 2026Brandon KaplanIndustry · Sep 8, 2026Clearstep Inc. (Bilal Naved, PhD, Co-Founder & Chief Product Officer)Industry · Sep 15, 2026Hari Prakash ChanumoluIndustry · Aug 18, 2026Navid FarrIndustry · Sep 8, 2026Newton’s TreeIndustry · Sep 3, 2026OneSource Solutions InternationalIndustry · Aug 28, 2026OrinyxIndustry · Sep 7, 2026Princeton Medical Systems (John Xavier, U.S. Partnerships & Regulatory Liaison)Industry · Sep 10, 2026Prof. Ray O'Sullivan (Vox / VoxMedical; Royal College of Surgeons Ireland)Industry · Sep 15, 2026Profound Ventures | Guidance Global Consulting (Brian Meshkin, Managing Partner; Anita Monteiro, CEO)Industry · Sep 14, 2026Ravi Pankhaniya, MDIndustry · Aug 28, 2026Sam Rosenthal (Red Kit)Industry · Sep 9, 2026Sentir Health, Inc. (Mario Ricart, Founder)Industry · Sep 12, 2026Shara GospelIndustry · Aug 24, 2026SichGate Inc.Industry · Aug 22, 2026Sitora Healthcare DigitalIndustry · Sep 3, 2026Steven Zhao (Independent Medical Device Regulatory Practitioner)Industry · Sep 14, 2026Supernova TechnologiesIndustry · Aug 18, 2026Tanmaya Kumar (Behavioral Health Open Source)Industry · Aug 26, 2026The Christman AI ProjectIndustry · Sep 4, 2026The Christman AI ProjectIndustry · Sep 4, 2026VivaSecurisIndustry · Aug 25, 2026Walnut Hill MedicalIndustry · Aug 18, 2026WhaleTeq Co., Ltd.Industry · Sep 8, 2026Yassen Eltayeb (Founder, Conefia LLC)Industry · Sep 12, 2026Chirag KanitkarClinicians · Aug 28, 2026Deborah Ault, RNClinicians · Aug 22, 2026Douglas Stoddard, MD (CHRISTUS Health)Clinicians · Aug 18, 2026Gregory Marcisz, CBETClinicians · Aug 24, 2026Manuj Agarwal, MDClinicians · Sep 3, 2026Michelle Bernabe, RN, BSNClinicians · Sep 10, 2026Sihem KhelifaClinicians · Sep 9, 2026Wen Hsien Ethan Huang, MDClinicians · Sep 3, 2026Joel GrunhutPublic / patients · Sep 7, 2026Qiong LiuPublic / patients · Sep 11, 2026Xiangyu Guo (Independent Researcher)Public / patients · Sep 13, 2026Martin HaimerlAcademia / other · Sep 1, 2026Rohith Reddy Bellibatlu (Independent Researcher, Clinical AI Evaluation Methodology)Academia / other · Sep 14, 2026Sehouenou Alberic Candide AhouehomeAcademia / other · Aug 29, 2026