Related Experiment Video
Updated: Jan 29, 2026

Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
Evaluating reasoning large language models with human-like thinking in ophthalmic question answering
Zhouqian Wang1, Chenjia Xu1, Lei Wang2
1Ningbo Key Laboratory of Medical Research on Blinding Eye Diseases, Ningbo Eye Institute, Ningbo Eye Hospital, Wenzhou Medical University, Ningbo, China.
Reasoning large language models (LLMs) show improved performance in ophthalmology question answering, with DeepSeek-R1 excelling. These models better simulate human thinking, paving the way for more reliable AI in eye care.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology
- Natural Language Processing
Background:
- Evaluating the capabilities of large language models (LLMs) in specialized medical domains like ophthalmology is crucial.
- Assessing the reasoning processes, not just accuracy, of LLMs is essential for trustworthy AI applications.
- Current LLMs may not fully replicate the nuanced thinking required for complex medical questions.
Purpose of the Study:
- To assess the performance of reasoning large language models (LLMs) in answering ophthalmology-related questions.
- To compare reasoning LLMs against conventional non-reasoning LLMs using a novel evaluation framework.
- To analyze the accuracy and reasoning quality of LLMs in ophthalmic question answering.
Main Methods:
- Evaluated two reasoning LLMs (DeepSeek-R1, QwQ-32B) and one non-reasoning LLM (LLaMA-3.3-70B-Instruct).
- Utilized MedQA-Eye, a custom dataset of 967 ophthalmology questions across 10 subspecialties and 3 languages.
- Developed a novel framework to assess LLM thinking patterns, simulating human medical reasoning.
Main Results:
- DeepSeek-R1 achieved the highest answer accuracy (90.59%), outperforming LLaMA-3.3-70B-Instruct (87.90%) and QwQ-32B (84.28%).
- Incorrect logical inference was the primary failure mode for reasoning LLMs (93.41%-94.74% of errors).
- DeepSeek-R1 demonstrated significantly lower semantic uncertainty (1.04±3.63) compared to QwQ-32B (4.31±40.70), indicating more reliable reasoning.
Conclusions:
- Reasoning LLMs, particularly DeepSeek-R1, show superior performance in ophthalmic question answering compared to non-reasoning models.
- The findings suggest reasoning LLMs can better mimic human-like thought processes in medical contexts.
- This advancement holds promise for developing more trustworthy and sophisticated AI systems in ophthalmology.
Related Concept Videos
Reason and Intuition
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Counterfactual Thinking
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...

