关于逻辑观察的大型语言模型的表现标识符名称和代码映射在实验室医学:对ChatGPT-4.0,Gemini和Perplexity进行比较分析
Shinae Yu1, Eun-Jung Cho2, Sollip Kim3
1Department of Laboratory Medicine, Haeundae Paik Hospital, Inje University College of Medicine, Busan, South Korea.
International journal of medical informatics
|January 8, 2026
概括
大型语言模型 (LLM) 显示了通过逻辑观察标识符名称和代码 (LOINC) 映射标准化实验室数据的潜力. ChatGPT-4.0和Perplexity AI的表现优于Gemini,但专家验证对于临床准确性至关重要.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 实验室医学 实验室医学
背景情况:
- 标准化医疗保健数据,特别是实验室测试结果,对于互操作性和准确的临床决策至关重要.
- 逻辑观察标识符名称和代码 (LOINC) 是用于识别实验室和临床观察的通用标准.
- 手动LOINC映射是耗时的,需要专门的专业知识.
研究的目的:
- 评估大型语言模型 (LLM) 在实验室医学中用于LOINC映射的可行性和实用性.
- 为了比较ChatGPT-4.0,Gemini 1.5和Perplexity AI在将实验室测试项目映射到LOINC代码中的性能.
主要方法:
- 选择了75个实验室测试项目 (55个临床化学,20个血液学).
- 六位临床病理学家建立了一个共识LOINC映射作为黄金标准.
- 将LLM输出 (ChatGPT-4.0,Gemini 1.5,Perplexity AI) 与黄金标准进行比较,将结果分类为完全匹配 (CM),部分匹配 (PM) 或不匹配 (MM).
主要成果:
- 在LLMs中观察到LOINC映射性能的显著差异.
- 聊天GPT-4.0和Perplexity AI显示了相似的性能,表现优于Gemini 1.5.5.
- 聊天GPT-4.0在临床化学CM (58.2%) 中领先,而Perplexity AI在血液学CM (55.0%) 中领先. 双子1.5具有最高的不匹配率 (80.0%在血液学).
结论:
- 简单的LLM可以帮助LOINC绘制,从而减少工作量.
- 结构化的输入,本地化和专家监督对于可靠的LLM生成映射至关重要.
- 人类验证对于确保LLM衍生LOINC代码的临床准确性至关重要.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
482
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Genetic Lingo
113.7K
Overview
113.7K
