An evaluation framework for clinical use of large language models in patient interaction tasks

Shreya Johri1, Jaehwan Jeong1,2, Benjamin A Tran3

  • 1Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.

Nature Medicine
|January 3, 2025
PubMed
Summary

This study introduces CRAFT-MD, a new method for testing clinical large language models (LLMs) through natural conversations. Current LLMs show limitations in diagnostic reasoning and accuracy, highlighting the need for better evaluation frameworks in medicine.

Related Concept Videos