基于大型语言模型的聊天机器人与临床医生的可靠性作为牙科信息来源:比较分析
Stefano Martina1, Davide Cannatà1, Teresa Paduano1
1Department of Medicine, Surgery and Dentistry "Scuola Medica Salernitana", University of Salerno, Via Allende, 84081 Baronissi, Italy.
Dentistry journal
|August 27, 2025
概括
大型语言模型 (LLM) 聊天机器人在回答正问题方面表现出很高的一致性,但往往与牙医有很大的不同. 虽然它们很有用,但也可能在复杂的话题上提供误导性的信息.
科学领域:
- 牙科科学
- 人工智能
背景情况:
- 大型语言模型 (LLM) 聊天机器人越来越多地用于信息检索.
- 它们在牙科等专业领域的可靠性需要彻底评估.
研究的目的:
- 评估LLM聊天机器人作为牙科信息来源的可靠性.
- 将聊天机器人的反应与一般牙医 (GDPs) 和牙科专家 (Os) 的反应进行比较.
主要方法:
- 我们向五个领先的聊天机器人提出了八个正确/错误的正确问题.
- 使用Cronbach的alpha测量聊天机器人响应的一致性.
- 通过基平方测试 (p < 0. 05) 将聊天机器人的反应与临床医生的反应进行比较.
- 通过全球质量表 (GQS) 评价教育价值.
主要成果:
- 所有聊天机器人显示出高响应的一致性 (alpha > 0.80).
- 在大多数问题上,聊天机器人和临床医生的反应之间观察到显著的差异 (p < 0. 05).
- DeepSeek获得最高的GQS评分 (中位数为4.00),而Copilot获得最低的评分 (中位数为2.00).
结论:
- 与牙科专业人员相比,LLM聊天机器人提供一致的,但往往不同的信息.
- 聊天机器人可以提供有价值的牙科见解, 但对于有争议的话题可能不可靠.
更多相关视频
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
681
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
575
相关概念视频
Modeling in Therapy
145
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
145
Stereotype Content Model
14.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.9K
