基于人工智能 (AI) 的三种大型语言模型在标准化测试中的表现;对人工智能辅助牙科教育的影响
Hamoun Sabri1,2, Muhammad H A Saleh1, Parham Hazrati1
1Department of Periodontics and Oral Medicine, School of Dentistry, University of Michigan, Ann Arbor, Michigan, USA.
Journal of periodontal research
|July 20, 2024
概括
与谷歌双子和ChatGPT-3.5.5相比,ChatGPT-4在回答美国牙周学会 (AAP) 考试问题方面表现出更高的准确性. 这突显了大型语言模型 (LLM) 作为周期学教育工具的潜力.
科学领域:
- 人工智能在牙科教育中的应用
- 牙周病学中的大语言模型 (LLM)
背景情况:
- 人工智能和自动化数据分析的整合将改变牙科教育.
- 这项研究旨在利用人工智能来增强教学教学方法.
研究的目的:
- 量化和比较ChatGPT (GPT-4,GPT-3.5) 和谷歌双子座对人类研究生响应的准确性.
- 评估LLM在美国牙周病学会 (AAP) 年度入职考试中的表现.
主要方法:
- 采用了一个比较的横截面研究设计.
- 从AAP在职考试 (2020-2023) 的1312个问题被管理到LLMs,并与牙周居民得分进行比较.
- 智方测试用于响应分析,对考试部分和难题进行分分析.
主要成果:
- 聊天GPT-4实现了最高的准确性 (平均79.57%),超过了人类对照,GPT-3.5和双子座 (p < .001).
- 谷歌双子 (平均72.86%) 超过了第一年和第二年居民,但没有超过第三年居民.
- 与GPT-4和Gemini相比,ChatGPT-3.5在所有考试年表现较差.
结论:
- 聊天GPT-4在回答AAP考试问题方面表现出显著的准确性和可靠性,表明其作为周期学教育工具的潜力.
- 双子座和ChatGPT-3.5表现较差,突出了需要改进的领域.
- 局限性包括无法处理基于图像的问题和响应不一致;需要进一步的研究.
相关概念视频
Language and Cognition
340
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
340
Learning Disabilities
93
Learning disabilities are cognitive disorders caused by neurological impairments that affect cognitive functions like language and reading, without indicating overall intellectual or developmental challenges. These disabilities differ from global intellectual or developmental disabilities as they are limited to distinct cognitive functions. Common learning disabilities include dysgraphia, dyslexia, and dyscalculia, each of which impacts unique aspects of learning.
Dyslexia
Dyslexia is a...
Dyslexia
Dyslexia is a...
93


