相关实验视频
Updated: Jan 18, 2026

05:40
Using an Automated Hirschberg Test App to Evaluate Ocular Alignment
Published on: March 24, 2020
15.7K
使用ChatGPT-4,微软Copilot和谷歌Gemini进行儿童眼科问题的比较
Tevfik Serhat Bahar1, Olgar Öcal2, Asli Çetinkaya Yaprak2
1The Department of Ophthalmology, Antalya Korkuteli State Hospital, Antalya, Turkey.
Journal of pediatric ophthalmology and strabismus
|May 27, 2025
概括
与ChatGPT和谷歌双子相比,微软Copilot在回答儿科眼科问题的准确度更高. 虽然人工智能聊天机器人提供潜在的好处,但由于其响应可能存在不准确性,用户应谨慎使用.
科学领域:
- 眼科医生 眼科 眼科
- 人工智能的人工智能
- 医学教育 医学教育
背景情况:
- 像ChatGPT,谷歌Gemini和微软Copilot这样的人工智能 (AI) 聊天机器人越来越多地用于信息检索.
- 这些人工智能工具在专业医疗领域的准确性和可靠性,如儿科眼科,需要进行彻底的评估.
- 评估AI性能对于理解它们在医学教育和实践中的潜在作用至关重要.
研究的目的:
- 为了评估ChatGPT,Google Gemini和微软Copilot在回答儿科眼科多项选择问题的准确性.
- 为了相互比较这三个领先的人工智能程序的性能.
- 在医学信息的背景下,评估人工智能生成的响应的可读性.
主要方法:
- 来自Ophtho-Questions在线银行的100个多选题被管理到ChatGPT,Gemini和Copilot.
- 人工智能生成的答案与官方答案密钥进行了比较,以确定正确性.
- 使用弗莱什-金凯德等级水平,弗莱什阅读易度得分和科尔曼-利乌指数分析了答案的可读性.
主要成果:
- 微软Copilot获得了最高的正确答案率 (74%),其次是ChatGPT (61%) 和谷歌双子 (60%).
- 副驾驶员的准确性在统计学上显著高于ChatGPT (P = .049) 和Gemini (P = .035).
- 可读性分析表明,Copilot的响应平均得分最高,而ChatGPT和Gemini的复杂性高于推水平.
结论:
- 与ChatGPT和谷歌双子相比,微软Copilot在回答儿科眼科问题的准确性更高.
- 人工智能聊天机器人可以成为儿童眼科知识获取的宝贵资源.
- 用户必须批判性地评估人工智能产生的内容,因为可能存在不准确性.
更多相关视频
11:12Driving Simulation in the Clinic: Testing Visual Exploratory Behavior in Daily Life Activities in Patients with Visual Field Defects
Published on: September 18, 2012
17.8K
05:32Comparing Eye-tracking Data of Children with High-functioning ASD, Comorbid ADHD, and of a Control Watching Social Videos
Published on: December 7, 2018
9.5K