在创建牙科板式问题时评估大型语言模型:一个前性的跨部分研究
Nguyen Viet Anh1, Nguyen Thi Trang1
1Faculty of Dentistry, Phenikaa University, Hanoi, Vietnam.
概括
这项研究比较了五种大型语言模型 (LLM) 用于生成牙科董事会式的多选择题 (MCQ). 克劳德3.5 索内特在提供答案的理由方面表现出色,展示了LLMs的LLMs.
科学领域:
- 牙科教育中的人工智能
- 用于医学问题的自然语言处理生成问题
背景情况:
- 对于牙科多选题 (MCQ) 生成的大型语言模型 (LLM) 的研究有限,先前的研究集中在ChatGPT和Gemini上.
- 这项研究解决了在创建牙科董事会式MCQ时对LLM绩效进行更广泛评估的需要.
研究的目的:
- 评估和比较五种先进的大型语言模型 (LLM) 在生成牙科板式问题的性能.
- 根据清晰度,相关性,适合性,分散注意力的质量和理由来评估生成的MCQ的质量.
主要方法:
- 五个LLM (ChatGPT-4o,Claude 3.5 Sonnet,Copilot Pro,Gemini 1.5 Pro,Mistral Large 2) 被用来从最近的牙科指南中生成350个MCQ.
- 两个独立的调查人员使用五个标准的10点利克特级别对每个问题进行了评估.
- 评估者之间的可靠性使用kappa评分进行评估.
主要成果:
- 取得了相当大的评价者间可靠性 (kappa:0.7-0.8).
- 所有的LLM都产生了清晰度,相关性和合理性 (9以上) 的高中等分数的问题.
- 克劳德 3.5 索内特在生成问题理性方面表现优越 (p < 0.01).
结论:
- 大型语言模型显示出产生高质量,临床相关的牙科板式问题的巨大潜力.
- 克劳德3.5 索内特 (Claude 3.5 Sonnet) 成为产生具有强大理由的MCQ的高性能LLM.
更多相关视频
相关概念视频
Assessment of the Mouth
416
A thorough mouth assessment, including inspection and palpation of the lips, gums, tongue, tonsils, uvula, and pharynx, is crucial in detecting potential health issues. Diseases ranging from oral cancer to systemic conditions like diabetes could be identified early through careful oral examination. This article provides a detailed guide on conducting a comprehensive mouth assessment.
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
416
Longitudinal Research
12.5K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.5K


