Large language models are less effective at clinical prediction tasks than locally trained machine learning models

Katherine E Brown1, Chao Yan1, Zhuohang Li2

  • 1Department of Biomedical Informatics, Vanderbilt University Medical Center (VUMC), Nashville, TN 37203, United States.

Summary

Traditional machine learning (ML) significantly outperforms large language models (LLMs) like GPT-3.5 and GPT-4 for clinical prediction tasks using electronic health records (EHRs). LLMs show limitations in performance, calibration, and privacy protection compared to traditional ML methods.

Related Concept Videos