评价人工智能产生的脊椎形教育材料的可读性和质量:对五种语言模型进行比较分析
Mengchu Zhao1, Mi Zhou2, Yexi Han3
1Department of Rehabilitation Medicine, Shanghai Sixth People's Hospital, Shanghai Jiaotong University School of Medicine, Shanghai, China.
Scientific reports
|October 10, 2025
概括
人工智能 (AI) 工具提供脊椎病信息,但质量各不相同. DeepSeek-R1提供了最易读的内容,尽管所有AI模型都缺乏引用,这影响了患者教育的可信度.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 脊柱健康教育 脊柱健康教育
背景情况:
- 脊椎形术语和治疗复杂性挑战了患者的理解.
- 患者越来越多地使用人工智能获取健康信息,引发了对内容质量的担忧.
- 人工智能产生的健康内容可能显示出可读性差和错误信息风险.
研究的目的:
- 评估人工智能生成的关于脊椎病的内容的可读性和信息质量.
- 为了比较五种不同的人工智能模型在生成与脊椎结节有关的信息方面的表现.
- 识别人工智能模型,为脊椎病患者提供可访问和可靠的教育材料.
主要方法:
- 五个人工智能模型 (ChatGPT-4o,ChatGPT-o1,ChatGPT-o3 mini-high,DeepSeek-V3,DeepSeek-R1) 进行了关于先天性,青少年异常症和神经肌肉结的查询.
- 使用弗莱什-金凯德等级水平 (FKGL) 和弗莱什-金凯德阅读易度 (FKRE) 评估可读性.
- 使用DISCERN评分评估内容质量,通过Intraclass关联系数 (ICC) 评估评级者之间的可靠性.
主要成果:
- DeepSeek-R1 显示出卓越的可读性,具有最低的 FKGL (6.2) 和最高的 FKRE (64.5).
- 聊天GPT-o1和聊天GPT-o3所需的大学水平阅读技能 (FKGL>12.0) 较小.
- 在所有模型中,DISCERN的分数一致 (~50.5/80),表明了相当的质量,但所有模型都缺乏引用.
结论:
- 人工智能生成的脊柱形脊椎病教育材料在可读性方面表现出显著的变化.
- DeepSeek-R1提供了最容易访问的内容,但由于没有引用,可信度受到限制.
- 未来的AI开发应该优先考虑增强可读性,并整合引用机制以提高准确性和可靠性.
相关概念视频
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Assessment of Airway, Skin Color, and Use of Accessory Muscles
1.6K
A thorough assessment of respiratory health is paramount in clinical settings to identify and manage respiratory distress and ensure adequate oxygenation. This article elaborates on the critical aspects of respiratory evaluation, including airway assessment, skin color examination, and the observation of accessory muscle use, which are integral to effectively diagnosing and managing patients with respiratory conditions.
Introduction
The initial evaluation of a patient's respiratory system...
Introduction
The initial evaluation of a patient's respiratory system...
1.6K
Health Literacy
5.2K
Health literacy is an individual's or a community's capacity to comprehend, receive, read, and use relevant healthcare information and services. The World Health Organization (WHO, 2018) defines health literacy as the cognitive and social skills that determine the ability of individuals to gain access to, understand, and use information in ways that promote and maintain good health. As a result, the WHO helps individuals manage long-term health concerns, participate in preventative...
5.2K
Nursing Evaluation
4.1K
The evaluation stage signals the end of the nursing process. The nurse gathers evaluative data to assess whether or not the patient has attained the expected results. Whereas the nurse collects data in the nursing assessment to identify the patient's health concerns, the evaluation stage data determines if the indicated health issues are resolved. Evaluative data collection includes two sections: the data acquired to evaluate patient outcomes and the time criteria for data collection.
4.1K
Models of Health Promotion and Illness Prevention I
2.7K
A model is a theoretical way to understand a concept or an idea. Models can overcome barriers to health regardless of diverse economic and cultural backgrounds. In addition, models make the task easier by providing different ways to approach complex issues. There are two major health promotion models: the health belief model and the health promotion model.
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
2.7K

