评估外吸收的提取增强的大型语言模型:双子座和笔记本的比较研究LMLM
Marc Garcia-Font1, Nicolás Dufey-Portilla2, Fernando Durán-Sindreu1
1Department of Endodontics, Universitat International de Catalunya, School of Dentistry, Barcelona, Spain.
Journal of endodontics
|November 9, 2025
概括
两个人工智能模型,谷歌双子和笔记本LM,在回答有关外部宫再吸收的临床问题时表现出高准确性和一致性. 笔记本LM的性能略有改善,但检索增强并没有显著改善这些任务的响应.
科学领域:
- 人工智能在牙科中的应用
- 临床决策支持系统 临床决策支持系统
- 在医疗保健中的自然语言处理.
背景情况:
- 这项研究评估了Alphabet Inc.两种大型语言模型的准确性和一致性:谷歌双子 (GG) 和笔记本LM (NLM).
- 评估的重点是回答与使用提取增强框架的外部宫再吸收相关的临床问题.
- 笔记本LM是一个基于文档的配置,而谷歌双子座则用于其基本配置.
研究的目的:
- 评估谷歌双子和笔记本LM在回答有关外部宫再吸收的临床问题的准确性和一致性.
- 为了比较一个基本的大型语言模型的性能与一个基于文档的配置.
- 确定检索增强是否对结构化临床任务的响应质量产生重大影响.
主要方法:
- 四十六个关于外部宫再吸收的二分组临床问题是由三个内牙专家创建的.
- 每个问题都通过三个独立的用户帐户向Google Gemini和NotebookLM提出,总共产生了276个答案.
- 答案由三个内牙专家独立评估,与准确性和一致性的黄金标准答案对比.
主要成果:
- 谷歌双子实现了89%的准确性和93%的一致性.
- 笔记本LM实现了96%的准确性和90%的一致性.
- 两种模型之间在准确性和一致性方面没有发现统计学上显著的差异.
结论:
- 在回答临床问题时,NotebookLM 和 Google Gemini 均表现出高准确性和一致性.
- 与谷歌Gemini相比,笔记本LM的性能略高.
- 检索增强对这些特定的结构化临床问题没有产生显著的改善.
更多相关视频
07:22Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
951
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
478
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
