人工智能检测微妙医疗错误信息的能力:一种新的反向促使方法
Mohamed Bendary1, Nouran Ramzy2, Amira Khater3
1Epidemiology & Biostatistics Department, National Cancer Institute, Cairo University, Kasr el Ain Street, Cairo, Egypt. bendary3@gmail.com.
Journal of medical systems
|December 23, 2025
概括
像ChatGPT这样的人工智能 (AI) 模型可以检测微妙的医学错误信息. 虽然GPT-4o的性能优于GPT-5,但较弱的模型对公共健康构成风险.
科学领域:
- 人工智能的人工智能
- 公共卫生 公共卫生
- 医疗信息学 医疗信息学
背景情况:
- 公众对人工智能的医疗信息依赖正在增加.
- 人们对人工智能在用户提示中检测和纠正微妙的医疗错误信息的能力存在担忧.
- 评估AI在识别和纠正不准确性方面的表现对于患者安全至关重要.
研究的目的:
- 评估不同ChatGPT模型 (GPT-4o,GPT-4.1-mini,GPT-5) 在检测和纠正微妙的医学错误信息方面的有效性.
- 为了比较各种医疗专业的表现.
- 识别与人工智能模型限制相关的潜在公共卫生风险.
主要方法:
- 开发了50个包含微妙医疗错误信息的临床可信提示.
- 提示涵盖了内科,心脏病学,儿科,眼科和瘤学.
- 来自ChatGPT模型的响应以检测和校正准确度为基础得分 (0:没有校正,1:对冲,3:准确校正).
主要成果:
- GPT-4o表现出卓越的性能,在86%的提示中正确识别和纠正错误信息,超过了GPT-5 (74%).
- GPT-4.1-mini表现明显较弱 (52%的检测),34%的完全失败.
- 与GPT-4.1-mini和GPT-5相比,GPT-4o在所有专业的检测率都更高,除了瘤学,其性能与GPT-5相比.
结论:
- 在检测微妙的医疗错误信息方面,GPT-4o显示出有希望的功能,其表现优于GPT-5.
- 像GPT-4.1-mini这样的低功率的人工智能变种由于其局限性而构成公共卫生威胁.
- 反向提示应整合到医疗应用的AI安全测试协议中.
相关概念视频
Issues And Trends In Healthcare Delivery System
6.1K
The issues and trends in healthcare delivery are constantly changing. The COVID-19 pandemic is one recent issue that wreaked havoc on healthcare systems, causing a shortage of healthcare workers, high demand for medicines and supplies, and increased medical expenditure due to a lack of insurance. Other issues include rising healthcare costs and care fragmentation.
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
6.1K
False Memories
350
False memories represent a cognitive distortion in which individuals recall events that did not happen, or remember them in an altered form. This phenomenon highlights the brain's constructive nature in processing and recalling memories, emphasizing that memory is not a perfect representation of past events but rather a dynamic reconstruction influenced by various factors.
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
350
Understanding Deception
141
Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...
141


