大規模言語モデルの肺がん臨床意思決定におけるパフォーマンス:DeepSeek、Grok、GPTに基づく比較分析
Yuyang Zhang1, Dandan Yang2, Yifan Shi3
1Department of Surgery, Jinzhou Medical University, Jinzhou, CHN.
Abstract:
Large language models (LLMs) have reached a breakthrough in many aspects of imaging analysis and guideline mining, but not enough research has been conducted on applying them specifically to lung cancer applications. Three models were chosen, and corresponding questions that highlighted specificity towards the diagnosis of lung cancer were proposed with the goal of providing data to increase confidence and improve recommendations for transforming AI-driven clinical care for lung cancer. In this study, five clinical domains were defined. Each question was individually uploaded to the models, and responses were evaluated by three thoracic surgery experts based on accuracy, completeness, and practicality. DeepSeek-R1, Grok-3, and GPT-4.5 showed different levels of results when it came to providing clinical support for lung cancer. Regarding their responses to clinical questions, Grok-3 had a much longer average response and better performance scores. Subgroup analyses further showed that Grok-3 scored the highest of all five domains. Additionally, the confidence scores on recognizing images in the text of Grok-3 were the highest; most mistakes occurred in the differential diagnosis of special cases, whereas DeepSeek and GPT all gave preference to rarer or infectious diseases.

