在双语复杂眼科推理中,DeepSeek-R1的表现优于Gemini 2.0 Pro,OpenAI o1和o3-mini
Pusheng Xu1, Yue Wu1, Kai Jin2
1School of Optometry, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China.
Advances in ophthalmology practice and research
|July 18, 2025
概括
在复杂的眼科问题上,DeepSeek-R1的表现优于其他大型语言模型 (LLM). 本研究评估了双语医疗场景中的LLM准确性和推理.
科学领域:
- 人工智能在医学中的应用
- 眼科医生 眼科 眼科
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在医疗保健中越来越多地使用.
- 评估他们在眼科等专业医疗领域的表现至关重要.
- 双语能力对于全球医疗应用是必不可少的.
研究的目的:
- 与其他LLM相比,评估DeepSeek-R1的准确性和推理能力.
- 为了评估复杂的双语眼科病例问题的表现.
- 在这个领域内识别LLM中常见的推理错误.
主要方法:
- 使用了来自中国眼科考试 (诊断和管理) 的130个多选择题 (MCQ).
- MCQs被翻译成英语,以创建一个双语测试集.
- 对DeepSeek-R1,Gemini 2.0 Pro,OpenAI o1和OpenAI o3-mini的响应进行了分析,以确定其准确性和逻辑推理.
主要成果:
- DeepSeek-R1获得了最高的准确性:0.862在中文和0.808在英语MCQs.
- 深度搜索-R1显著优于其他模型 (P <0.001的中文,P <0.05的英语).
- 常见的推理错误包括忽视关键患者病史/症状以及误解医疗数据.
结论:
- 在双语眼科推理任务中,DeepSeek-R1表现出卓越的性能.
- 先进的LLM显示了协助临床决策的潜力.
- 建议在医学背景下评估LLM推理的框架.
相关概念视频
Blind Procedures
12.1K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
12.1K
Reasoning
132
Reasoning is the action of thinking about something in a logical, sensible way. It is integral to problem-solving, decision-making, and critical thinking. Reasoning can be inductive or deductive. Reasoning involves transforming information into conclusions, which is essential for problem-solving, decision-making, and critical thinking.
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
132
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Machines: Problem Solving II
373
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
373


