评估大型语言模型作为医学短答题的评分器:与专家人类评分器进行比较分析
Olena Bolgova1, Paul Ganguly1, Muhammad Faisal Ikram1
1College of Medicine, Alfaisal University, Riyadh, Kingdom of Saudi Arabia.
Medical education online
|August 24, 2025
概括
大型语言模型 (LLM) 显示出对医学短答题的评分有希望,但性能因模型和医学学科而异. 专家标准并没有持续提高LLM的准确性,因此需要特定领域的实施和人类监督.
科学领域:
- 医学教育评估
- 医疗保健中的人工智能
- 自然语言处理
背景情况:
- 医学教育对简短答案问题的评估是耗时的,需要专家的意见.
- 大型语言模型 (LLM) 是自动化SAQ分级的潜在解决方案.
- 在专业医疗评估环境中,LLM的有效性尚未得到充分证实.
研究的目的:
- 评估五个LLM对医疗SAQ的人类评分专家的评分表现.
- 在不同医学学科 (解剖学,组织学,胚胎学,生理学) 中比较LLM能力.
- 评估提供专家定义的评分标题对LLM绩效的影响.
主要方法:
- 在四个医学科目中分析了804名学生的SAQ反应.
- 通过三个专家的人类教师对答案进行评分.
- 使用两种方法评估五个LLM (GPT-4.1,Gemini,Claude,Copilot,DeepSeek):自行生成的标准和专家提供的标题.
- 使用科恩的卡帕和类内相关系数 (ICC) 来衡量协议.
主要成果:
- 观察到实质性的专家-专家协议 (平均卡帕:0. 69,ICC:0. 86).
- 在不同类型和模型中,LLM的成绩差异很大,没有一个LLM的成绩始终优于其他.
- 克劳德 (卡帕:0.61) 在一个问题上和DeepSeek (卡帕:0.53) 在另一个问题上达成了最高的专家-LLM协议.
- 专家标准对LLM业绩的影响不一致.
- 与专家相比,LLM在分级严格度上表现出显著的差异,并显示出特定领域的分级差异.
结论:
- 虽然LLM具有支持医疗SAQ评估的潜力,但具有特定领域的性能差异.
- 提供专家标题并没有持续提高LLM分级的准确性.
- 在医学教育评估中成功实施LLM需要仔细考虑特定领域和持续的人类监督.
更多相关视频
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
681
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
575
相关概念视频
Classification of Illness
9.4K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
9.4K
Sensitivity, Specificity, and Predicted Value
1.9K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.9K
