Check that evidence fits the intended users and setting
Proposes empirical testing of human oversight in representative workflows, without a clear prospective-study threshold.
Read the source passage
2. The risk framework: “human oversight” must be tested, not assumed Response to Discussion Questions 1, 2, and 14 The two-axis framework (device activity × severity of harm) is sound. Question 1 asks whether additional dimensions — including the time pressure of the deployment setting — should be represented. My answer is yes, and specifically: the degree of human oversight should be treated as an empirical property of the deployment setting, not as a design feature that is present or absent. In teaching clinicians to work with AI, the hardest lesson is this: the presence of an override option does not guarantee the override will be used. Automation bias is well documented, and clinicians under time pressure defer to confident outputs. Appendix A element E.4 recognizes automation bias, but treats it as a communication-quality attribute of the device. I would encourage CDRH to also treat it as a modifier of position on the activity axis: a function nominally placed at “acts with continuous HCP supervision” may in practice operate closer to autonomy if the supervision is not exercised. I encourage FDA to: Treat “degree of human oversight” as a property to be demonstrated in representative use conditions — time-pressured, multi-patient, real interface — rather than asserted in labeling. Ask sponsors to show evidence that intended users can and do detect incorrect outputs in representative workflows. This is an override-rate and detection-rate measurement, and it is precisely the kind of human-AI team evidence contemplated in Question 14. Note that the paper’s own observation — that a “talk to your doctor” statement may not make an output less directive — applies symmetrically, and bears on Question 2: an override interface that is never used provides no oversight. Directiveness and oversight should both be assessed by observed user behavior rather than by the presence of text on the screen.Original source ↗