Judge against the applicable standard of care · Use clinicians matched to the clinical task · Evaluate the clinician and AI working together
Counts the explicit approaches or boundaries identified in this passage. Categories can overlap; the stated clinical scope still applies.
Read the source passage
R7. The comparator should be the standard of care, not the "median clinician in practice" (Questions 14 and 15) This follows from S4 and S5. The paper offers two comparators: a clinician panel whose consensus reflects the standard of care, or the performance of a median clinician in practice. These are not equivalent. The median clinician in practice is, by definition, below the standard of care some of the time, and a comparator set at that level normalizes existing gaps in care as an acceptable ceiling for a new technology. A patient-first framework should set the standard of care as the floor, with prespecified non-inferiority margins justified by clinical context. Where a device is intended to extend specialist knowledge to generalists (a benefit the paper rightly names), the comparator for the device should be specialist-level performance on the specialist question, not generalist performance. Evaluation of the human-AI team should include measurement of automation bias — whether clinicians over-accept outputs — and not only measurement of the team's aggregate accuracy, since the former predicts how the team will behave when the device is wrong.Original source ↗