Related Experiment Videos
Evaluating Large Reasoning Models Versus Human Multidisciplinary Teams in Lung Cancer Decision-Making: Real-World
Ivan Viculin1, Josip Vrdoljak2, Krešimir Tomić3
1Department for Pulmonary Disease, University Hospital of Split, Split, Split-Dalmatia, Croatia.
Journal of Medical Internet Research
|July 16, 2026
Summary
Large reasoning models (LRMs) like GPT-5-Thinking show high-quality recommendations in lung cancer care, outperforming other models and human multidisciplinary teams (MDTs). Awareness of AI comparison did not alter MDT decision quality in this study.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Oncology
Background:
- Real-world evaluations of large language models (LLMs) and large reasoning models (LRMs) in clinical workflows are limited.
- Lung cancer care, relying on multidisciplinary team (MDT) integration, presents a complex environment for assessing LRMs.
Purpose of the Study:
- To compare the recommendation quality of two LRMs (GPT-5-Thinking and Deepseek-v3-r1) against each other and human MDT decisions in lung cancer cases.
- To evaluate the impact of MDT awareness of AI benchmarking on the quality of their decisions.
Main Methods:
- A single-center, real-world comparative study involving 100 lung cancer MDT cases.
- Deidentified case reports were submitted to GPT-5-Thinking and Deepseek-v3-r1 for diagnostic and therapeutic recommendations.
- Recommendations and MDT decisions were graded by two independent lung oncologists using Likert scales.
Main Results:
- GPT-5-Thinking consistently outperformed Deepseek-v3-r1 across diagnostic and therapeutic recommendations.
- Both LRMs generated recommendations that exceeded expert-graded MDT decisions in quality.
- MDT decision quality remained unaffected by awareness of AI benchmarking.
Conclusions:
- LRMs can produce high-quality recommendations in lung cancer care, with GPT-5-Thinking demonstrating superior performance.
- The integration of LRMs into MDT workflows may enhance clinical decision-making.
- Further prospective studies are needed to confirm the clinical utility of LRMs in improving patient outcomes.