Clinician Artifact Literacy Training for Safer Use of Medical Vision-Language Models During Diagnostic Imaging Education Sessions
Abstract
Medical vision-language models are increasingly shown to students, residents, and clinicians as tools for image explanation, case discussion, and preliminary educational interpretation. Their outputs can be persuasive even when the image contains acquisition artifacts, overlays, degraded contrast, or nonclinical visual marks. This creates an educational risk: learners may absorb the model’s explanation without recognizing that the answer has been shaped by image conditions rather than by reliable diagnostic evidence. This paper reports a controlled educational study of artifact literacy training for medical vision-language model use in diagnostic imaging education. We enrolled 186 participants across medical students, radiology residents, emergency medicine residents, and physician assistants. Participants reviewed model-assisted imaging cases before and after a structured training module focused on recognizing artifact-sensitive model behavior, distinguishing image evidence from model narrative, and documenting uncertainty. The intervention used chest radiography, musculoskeletal radiography, dermatology photographs, and ultrasound frames with benign artifacts, clinically relevant artifacts, and model-salient nonclinical marks. Compared with the control group, trained participants improved their artifact-related error detection from 41.8% to 73.6%, reduced inappropriate acceptance of model explanations by 28.9 percentage points, and wrote more specific uncertainty notes. The largest gains appeared among medical students and non-radiology trainees. Follow-up testing four weeks later showed partial retention, with artifact-related error detection remaining 19.4 percentage points above baseline. The findings suggest that safe educational use of medical vision-language systems requires explicit training in artifact literacy rather than simple warnings about model limitations