Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks

Eun Jeong Gong1,2,3, Chang Seok Bang1,2,3, Jae Jun Lee3,4

  • 1Department of Internal Medicine, Hallym University College of Medicine, Chuncheon, Gangwon, Republic of Korea.

PubMed
Summary

Large language models (LLMs) in medicine excel on exams but struggle in clinical practice. High scores do not guarantee patient safety, necessitating practice-based validation before deployment.

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K