Related Experiment Videos
Beyond the Black Box: A Point-Of-Care Framework for LLM Literacy and Critical Appraisal
Christopher R Stephenson1, Christopher A Aakre1, Ivana T Croghan1,2
1Department of Medicine, Division of General Internal Medicine, Mayo Clinic, Rochester, MN, USA.
Abstract:
Large language models (LLMs) are increasingly used in primary care for documentation, patient communication, and clinical decision support. However, many clinicians are adopting these tools for patient care without structured training or a clear understanding of their limitations. LLM output may appear authoritative but can be inaccurate, biased, and unsupported by evidence. This creates a growing gap in artificial intelligence LLM literacy. Clinicians need practical strategies to evaluate when LLM-generated information can be trusted and how to apply it to patient care. This article proposes an evidence-based framework for critically appraising the output of LLMs at the point of care. Building on principles from evidence-based medicine (EBM), we describe three levels of appraisal: 1) internal validation, 2) external verification, and 3) contextual application. Internal validation assesses whether the LLM output is stable, logically sound, and appropriately addresses uncertainty. External verification determines whether the output is accurate and supported by external evidence, including guidelines, peer-reviewed literature, or trusted clinical references. Contextual application evaluates whether the output is appropriate for the individual patient, taking into account comorbidities, social context, health literacy, patient preferences, and shared decision-making. The rigor of appraisal should be scaled to the clinical risk. Lower-risk uses, such as drafting patient education materials or simple communications, may require only editorial review. Higher-risk uses, such as diagnostic or management reasoning, require more rigorous appraisal before recommendations can be applied. LLMs may reduce cognitive load and improve efficiency, but they do not replace clinical judgment. Safe AI use in primary care depends not only on the sophistication of the LLM but also on the clinician's ability to critically evaluate, verify, and contextualize its output.