Related Experiment Videos

Performance of large language models in electrocardiogram interpretation: A comparative study

Gregory W Chai1, Samuel J Y Chen2, Jasmine Yang3

  • 1Temerty Faculty of Medicine, University of Toronto, Toronto, Ontario, Canada; Division of Cardiology, Keenan Research Center for Biomedical Science, St. Michael's Hospital, Unity Health Toronto, University of Toronto, Toronto, ON, Canada.

Summary

Large language models (LLMs) like ChatGPT and Gemini show limited accuracy in interpreting electrocardiograms (ECGs) for electrical axis and heart rhythm. Their current performance indicates they are not yet reliable for independent clinical diagnosis in cardiology.