提高人类现象型本体学识别的准确性:多模式大语言模型的比较评估
Wei Zhong1, Mingyue Sun2, Shun Yao3
1Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University, Beijing Maternal and Child Health Care Hospital, 251 Yaojiayuan Road, Chaoyang District, Beijing, 100020, China, 86 15572779093.
Journal of medical Internet research
|June 2, 2025
概括
多模式大语言模型 (MLLMs) 显著提高了初级医生在识别罕见疾病的人类表现型本体学 (HPO) 术语的准确性. 尽管幻觉率很高,但MLLM对罕见疾病诊断和表型标准化充满希望.
科学领域:
- 医疗信息学 医疗信息学
- 人工智能在医学中的应用
- 罕见疾病的诊断 罕见疾病的诊断
背景情况:
- 准确识别人类现象型本体学 (HPO) 术语对于罕见疾病的诊断和管理至关重要.
- 由于HPO的复杂性和手动搜索的局限性,初级医生在精确的表型描述方面面临挑战.
- 传统的HPO数据库搜索耗时且容易出现错误.
研究的目的:
- 评估多式大型语言模型 (MLLMs) 在提高初级医生从患者图像中识别HPO术语的准确性方面的有效性.
- 将MLLM辅助的HPO识别与传统的手动搜索方法进行比较.
主要方法:
- 10个专业的20名初级医生评估了27名罕见病患者的图像.
- 形成了两个组:使用中国HPO网站进行手动搜索,使用ChatGPT-4o提示程序进行MLLM辅助搜索.
- 准确性与专家定义的标准集相比进行了测量;记录了MLLM幻觉率.
主要成果:
- 通过MLLM辅助的组获得了67.4%的准确性,明显超过了手动组的20.4%的准确性 (P<.001).
- 独立的MLLMs (ChatGPT-4o,Llama3.2) 显示了可变的准确性 (15%-48%),但高的幻觉率 (不正确的ID和捏造的术语).
- 对罕见和遗传性疾病的医生培训可能会对HPO识别性能产生积极影响.
结论:
- MLLMs显著提高了初级医生的HPO术语识别准确度,有助于罕见疾病诊断和表型标准化.
- 在MLLM中显著的幻觉率需要在临床实施之前进行进一步的开发和验证.
相关概念视频
Background and Environment Affect Phenotype
6.5K
Although the genetic makeup of an organism plays a major role in determining the phenotype, there are also several environmental factors, such as temperature, oxygen availability, presence of mutagens, that can alter an organism’s phenotype.
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
6.5K
Improving Translational Accuracy
9.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.3K


