Shannon Kamalaker
What they argued
Q11 'don't cut corners, do a clinical trial'; humans must remain in place; likes competency approach but demands trial plus consent.
Themes it raises
FDA questions it names
Q1 · The two-axis risk frameworkQ2 · The spectrum of device activityQ3 · When an output becomes directiveQ4 · Generalist and specialist usersQ5 · Multi-turn conversations that migrateQ6 · Care escalation functionsQ7 · The competency-based approachQ8 · Mapping the risk grid to evidenceQ9 · The benchmarking structureQ10 · Benchmark contamination and saturationQ11 · Clinical confirmation without a prospective trialQ12 · Statistically meaningful performanceQ13 · Synthetic dataQ14 · Comparators and acceptance criteriaQ15 · Performance against usual careQ16 · Independent third partiesQ17 · Devices with many functions
Across the five cross-cutting questions
High-consequence work: Informs
The comment as filed
Here are my responses to this with the added disclaimer that I answered all questions to the best of my own brain ability and did not use AI in any form or fashion. Sincerely Dr. Shannon Nadia Kamalaker
Attachment
1. Does the two-axis risk framework, organized around device activity and the
consequence of relying on an incorrect output, appropriately capture the
dimensions most relevant to the risk of a GenAI-enabled software function? If
there are additional dimensions—such as the reversibility of a resulting action,
the availability of downstream safeguards, the time pressure of the deployment
setting, or the traceability of the output (i.e., to primary source materials)—that
should be represented in a risk framework, please describe and provide
examples of how they should be represented.
Response:
It does not appropriately capture the risk related to the Gen-AI enabled software
function.As a resident physician in the psychiatry space, it is imperative that we
are able to define clear and objective methods for the risk of the anything medical
advice made. Clearly we have always had several cases of issues and court
cases where the generativeAI is indicated and gave false information. For a
action-directing function a clear message involving ANY kind of medical
knowledge or guidance needs to be put and not just for physicians, but for
therapists, counselors, hotline specialists, etc. Patient or people need to
understand and realize that it is imperative and important to continue to use US
in the mental health field versus an AI model. One the model is limited in the
space of what actually can be said and given, and has to be trained using patient
data or physician data. To maintain the privacy and ethics of healthcareeverything would be consented too. Unsure how the generativeAI is done and
since it has already has some much data that is used, it would be impossible to
train the AI to this criteria without some ethical implications. Even
OpenevidenceAI. Also for the action-decision making methods. Again THIS is not
an person, researcher, physician or anything. There is a reason that many of us
go through medical school, residency, and research and beyond as well
continuing medical education.
2. CDRH seeks input on how to account for the spectrum of GenAI-enabled
informational functions that vary in the degree to which they direct a user to a
particular action (i.e., between “non-directive” and “action-directing”). What
characteristics of an output—such as its wording, specificity, personalization, or
context—could be considered as modifiers of the risk associated with an
informational function, after accounting for the device’s overall functionality and
intended use? What additional information would provide manufacturers with
sufficient clarity and predictability around risk assessment for informational
functions while recognizing that directiveness may exist along a continuum
rather than as a binary distinction?
Response:
For a action-directing function a clear message involving ANY kind of medical
knowledge or guidance needs to be put and not just for physicians, but for therapists,
counselors, hotline specialists, etc. Patient or people need to understand and realize
that it is imperative and important to continue to use US in the mental health field versus
an AI model. One the model is limited in the space of what actually can be said and
given, and has to be trained using patient data or physician data. To maintain the
privacy and ethics of healthcare- everything would be consented too. Unsure how the
generativeAI is done and since it has already has some much data that is used, it would
be impossible to train the AI to this criteria without some ethical implications. Even
OpenevidenceAI. Also for the action-decision making methods.
In terms of the additional information- every drug administration, every health clinic, hell
everyone that uses the drug would need to have a disclaimer that makes note that
generative AI or AI in general is not a fail safe and that your results may vary. Anything
and everything would need to be said with the behavioural health space that appropriate
mental health facilities and clinics needed to be conducted to ensure PATIENTS that
their safety is represented.
3. CDRH seeks input on whether and when a GenAI-enabled function that results
in delivery of clinical information to patients, as opposed to HCPs, could
present different or higher risks, while also recognizing the potential benefits
associated with improved patient empowerment, engagement, and access to
clinical information. What device characteristics, output features, or safeguards
might mitigate risks that could arise when a user lacks the domain knowledge
to independently evaluate an output, without unnecessarily underestimating
patient capability?
Response: For a action-directing function a clear message involving ANY kind of
medical knowledge or guidance needs to be put and not just for physicians, but for
therapists, counselors, hotline specialists, etc. Patient or people need to understand and
realize that it is imperative and important to continue to use US in the mental health field
versus an AI model. One the model is limited in the space of what actually can be said
and given, and has to be trained using patient data or physician data. To maintain the
privacy and ethics of healthcare- everything would be consented too. Unsure how the
generativeAI is done and since it has already has some much data that is used, it would
be impossible to train the AI to this criteria without some ethical implications. Even
OpenevidenceAI. Also for the action-decision making methods.
In exposing on this question, in every form that the AI is provided- we would need to
have it translated into different languages and even broken down to the form.
4. CDRH seeks input on whether and how the distinction between generalist and
specialist physicians might inform the assessment of risk for HCP-facing
GenAI-enabled functions. Under what circumstances, if any, might risk be
affected when an HCP who lacks the relevant clinical specialist knowledge
receives information that falls within a specialist area of practice? What device
characteristics or safeguards might mitigate such risks?
Response: Truthfully at the end of the day, with the idea of AI seemed really effective
and helpful, it is more work than anything. The same way that automated phone system
drive everyone crazy, so does AI. In terms of audits, mental health space, I have no
idea how this would be implemented. Surely from the energy standpoint, I DO NOT
want this implemented and neither do patients of all age size and race. In context, I
have stopped using any form of AI, have not even upgraded my gaming computer and
choose to put my stock into an IRA than to continue to invest in companies that support
AI. Nivdia and Microsoft I am looking at it. Again a clear warning that AI does not protect
privacy because nothing on the internet is protected, and clear implication that AI is
being used. No safeguard currently can migrate risk and the only way to fix these risks
would to still have humans in place
5. For multi-turn conversational GenAI-enabled devices that may migrate from
providing “non-directive” information to “action-directing” information over the
course of an exchange, how should risk be assessed across realistic
conversational trajectories? How could the intended use of such a device be
characterized when its behavior is emergent across a conversation?
Response:
The model that currently seems to work for many clinicians with a valid NPI number is
OpenEvidence AI. While many of us use it to help write prior authorizations and look up
new and updating criteria, the idea of physicians using it for patients privacy, messages,
etc that is now on the site is RIDICULOUS and frankly an insult to my intelligence. Over
time, it has improved limitations of certain questions, reading the message that
Openevidence AI that the question is out of the scope, but who is not the stop other
physician or medical students to use it elsewhere. In the end, I have just stopped using
it, mainly for that reason and because as a physician, I hold the same level of respect
and reasonability to patient and other colleagues- I write my own note and responses
across here..
6. For GenAI-enabled care escalation functions, CDRH is considering how the
evaluation may account for both under-escalation and over-escalation. How
could manufacturers characterize and weigh these two directions of error, given
that they may not be commensurable and that acceptable trade-offs may vary
by clinical context?
Response: Manufactures can put their disclaimer, but in the end, anyone that
participates in the use of AI in this space bares reasonability. The government system
itself, no one is immune to the consequences that will come with it. There are no
acceptable trade-offs by a clinical context. As a physician the oath comes into play: do
no harm, and since right now AI carries more risk and more reasonability, and frankly
I’m not training anything while not getting paid for, as well the ethical and environmental
impact, AI can kick rocks. I’d rather keep the medical degree that I paid for and invested
all my time and money in.
7. Is the competency-based approach described above, i.e., device benchmarking
followed by clinical confirmation, a useful and appropriate framework for
evaluating GenAI-enabled devices?
Response: The approach seems ideal in hindsight, but how will one maintain patient
safety and data during this? Who is not to say that the data gets leaked anyway, and in
that case who is responsible?
8. How might the two-axis risk framework described in Section IV be considered
within the competency-based approach described above to help determine the
level of evidence needed for a premarket submission?
Response: The two-axis framework is not even needed. Currently HIPPA laws and the
Security Rule of Healthcare govern that all healthcare providers within the space must
follow the same privacy and security that all would want--- how will the FDA,
government, manufactures, and physicians and other healthcare colleagues ensure that
patient data will be protected and not breached or stolen or sold?
9. Would a benchmarking structure such as the one described above be likely to
provide adequate evidence of clinical knowledge, analytic capabilities, safety
behavior, communication, and generalizability to support a reasonable
assurance of safety and effectiveness? Are there elements that are missing,
redundant, or inappropriately categorized? Are there externally developed
standards that could be leveraged?
Response: The benchmarking structure is nice, but a clinical trial or study would need to
be done and the idea of informed consent would have to be discussed. In essence as
well how does that maintain that patient confidentiality will be paramount while
delivering accurate and clinically appropriate data/info?
10. CDRH recognizes that publicly available benchmarking assets may be subject
to data contamination, saturation, and limited real-world representativeness.
How should a sponsor establish that performance on a given benchmark
predicts safe and effective real-world behavior for the device’s intended use?
What evidence should support the construct validity of a benchmark used to
gate device evaluation, and what role should sponsor-developed benchmarks
play given potential concerns around independence and optimization to the
test?
Response: No idea- again this benchmark really needs more fleshing out and there
would need to be a pilot study with the idea clearly laid out to be done. Also a
behavioral health study- our field is much more nuance and have more of a greater idea
and consequence if something were to go wrong. Again in terms of patient safety and
informed consent- who is responsible? These are federal laws and with them come
federal consequences.
11. CDRH is considering that clinical confirmation for a GenAI-enabled device
might not require a prospective clinical study in every case, and has described
above a range of approaches of increasing rigor and patient exposure. How
might a sponsor select and justify a confirmation approach tailored to a device’s
intended use and proportionate to the device’s risk profile? Are there device
types or risk profiles for which one or more of these approaches would be
insufficient or inappropriate? How might the anticipated distribution of real-world
inputs be taken into consideration? Are there other methods of clinical
confirmation that might help inform the evaluation of GenAI-enabled devices?
Response: No, don’t cut corners- do a clinical trial study and make it clear to everyone
including the public. THAT is not ok.
12. How should sponsors achieve statistically meaningful performance
measurement for GenAI-enabled devices? Where synthetically generated
inputs supplement real patient data, how should sponsors account for
differences between the synthetic and real-world distributions when estimating
performance, and under what conditions, if any, is it appropriate to combine
benchmarking evidence and clinical confirmation evidence to support a single
performance estimate?
Response: Qualtrics studies, accuracy of info given, patient info- but again dicey with
the informed consent
13. For which clinical domains, device functions, or subpopulations is synthetic
data particularly well-suited, or particularly inadequate, as a supplement to realworld evidence? What safeguards would mitigate the risk that synthetic data
generated by models of the same class as the device under evaluation
reproduces the very performance gaps the evaluation is intended to detect,
particularly for underrepresented subgroups?
Response: Almost everyone. Patients are the most likely to not understand something.
Behavioral health, elderly, children, patients in different languages and socioeconomic
factors may not be able to participate in this study well, rural populations. No idea how
one would be inclusive for all but If the study is limited, the data is limited in my position.
14. For open-ended device outputs, how should performance comparators and
acceptance criteria be selected? When a panel of qualified clinicians serves as
the comparator, how should the applicable standard (for example, the standard
of care versus the performance of a median clinician in practice) be defined and
justified? Should generalist or specialist physicians be used as a performance
standard? When should human-AI team performance, rather than the device
operating alone, serve as the basis for evaluation?
Response: I don’t believe this would make sense either, and again don’t have a answer
here. There would need to be qualified clinicians that would have to agree – maybe it
could be done with the hospital- but despite many physicians using AI I don’t know
many that would agree to actually participating in it for fear of the backlash, the society
and again the safety of themselves and their only license.
15. Are there ways in which performance might be assessed relative to the care,
technology, or course of action likely to occur in the absence of the device,
rather than to the comparators described in this section? What approaches
might be used to identify and justify a comparator such as unaided clinical
judgment, delayed specialist review, or no intervention?
Response: Not really answerable at this time, we would need a comparative device and
a pilot study to really have something to compare it too.
No unaided clinical judgement, like why bother doing the study at all. Delayed specialist
review would only delay care, which is dumb. No intervention must be an appropriate
answer but since it is due by a robot and not a person- who is say that it is justified? Are
the AI models going through medical school and residency and doing the same patient
care we are? Seeing the same patient populations, etc?
16. Should independent third parties be involved in some or all aspects of a
competency-based approach, including device benchmarking and clinical
confirmation? If so, in what ways might qualified, independent third-party
participation contribute to this approach, and what qualifications and
independence criteria should apply? Are there aspects of the assessment for
which third-party involvement would be impractical or inadvisable? What
safeguards or program design features would be critical to prevent third-party
participation from limiting competition or preventing innovation?
Response: An independent third-party would be create. All of the research boards, the
national institute of health, the mental health organizations like NAMI and SAMHSA,
AMA everyone should be involved. You wouldn’t get accurate results or information or
even support otherwise. I don’t think that third-party involvements of these above
organizations would be possible. Limiting design features is again privacy and data
protection.
17. CDRH is exploring whether a competency-based approach could be applied to
devices that incorporate a variety of underlying model architectures, such as
multimodal vision-language models and generative or predictive world models.
Are there aspects of the described approach that might be ineffective or
inapplicable to these devices? What additional or distinct considerations might
need to be addressed?
Response; I like the competency based approach but first yall need to flesh out the
public’s opinion on this in addition to generative AI- which it seems wrote these
responses. Reaching out to several different groups with these questions as well
involving the public’s opinion on this with a research paper that is written in more similar
terms would be ideal.
Disclaimer: These responses were written without the use of Generative AI. I do not
agree to this being published or edit or used without my written consent and permission.
I do not agree that this should be used in helping Generative AI define metrics. While I
recognized that everyone has a right their own opinion, I do not have to support it. Use
AI if you wish. I thought about posting this anonymously but I feel so strongly about this
topic, its now public. Since the government is now using AI with Medicare and Medicaid
audits, I want to let everyone know some of us care about patients and rather than use
AI actually get google with -AI because we care.
Sincerely,
Dr. Shannon Nadia Kamalaker