Performance of ChatGPT-4o and Four Open-Source Large Language Models in Generating Diagnoses Based on China's Rare

Wei Zhong1, YiFan Liu1, Yan Liu1

  • 1Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, No. 251 Yaojiayuan Road, Chaoyang District, Beijing, China, 8618810963279.

Summary

ChatGPT-4o achieved the highest diagnostic accuracy for rare diseases. Integrating retrieval augmented generation (RAG) significantly improved open-source large language models (LLMs), highlighting the need for careful model selection based on language and parameterization for clinical use.