评估由大型语言模型产生的眼膜病例的差异诊断
Jeffrey C Peterson1, Sruti S Rachapudi1, Sasha Hubschman1
1Department of Ophthalmology and Visual Sciences, University of Illinois at Chicago, Chicago, Illinois.
Ophthalmic plastic and reconstructive surgery
|August 11, 2025
概括
大型语言模型 (LLM) 显示出对眼膜成形诊断的前景. OcuSmart/EyeGPT和Claude 3.5在准确性和回忆方面表现出色,而Gemini在差异诊断方面提供了精确性.
科学领域:
- 眼科医生 眼科 眼科
- 人工智能的人工智能
- 医学诊断 医学诊断 医学诊断
背景情况:
- 准确的差异诊断对于有效的眼膜病患者护理至关重要.
- 评估新兴人工智能技术的诊断能力,如大语言模型 (LLM) 是必不可少的.
研究的目的:
- 评估六种不同的大型语言模型在生成眼膜状况的差异诊断中的准确性.
- 为了比较各种LLM与专家策划的诊断的性能.
主要方法:
- 来自EyeRounds.org的20个眼镜塑形病例被用来生成六个LLM的差异诊断:ChatGPT 3.5,ChatGPT 4.0,OcuSmart/EyeGPT,Google Gemini 1.5,Claude 3.5和微软CoPilot.
- 根据最高诊断匹配率,包括正确诊断,回忆和精度与专家差异相比,评估了LLM输出.
主要成果:
- OcuSmart/EyeGPT实现了最高的顶部诊断匹配率 (85%).
- 克劳德3.5显示了最高的正确诊断包括和回忆 (分别为100%和55%).
- 谷歌双子提供了最精确的差异 (43%),尽管克劳德3.5产生了不那么简洁的列表. 性能因案件类型而异.
结论:
- 临床眼膜显示出有助于诊断眼膜病例的巨大潜力.
- 在诊断和回忆方面,OcuSmart/EyeGPT和Claude 3.5表现强,而ChatGPT 3.5,OcuSmart/EyeGPT和Gemini则提供了简洁的差异.
- 需要进一步的验证和整合研究来将LLMs纳入临床实践.
相关概念视频
Prosopagnosia
249
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
249
Assessment of Airway, Skin Color, and Use of Accessory Muscles
1.1K
A thorough assessment of respiratory health is paramount in clinical settings to identify and manage respiratory distress and ensure adequate oxygenation. This article elaborates on the critical aspects of respiratory evaluation, including airway assessment, skin color examination, and the observation of accessory muscle use, which are integral to effectively diagnosing and managing patients with respiratory conditions.
Introduction
The initial evaluation of a patient's respiratory system...
Introduction
The initial evaluation of a patient's respiratory system...
1.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K


