Predicting Immunotherapy Response in Unresectable Hepatocellular Carcinoma: A Comparative Study of Large Language
Jun Xu1,2,3,4, Junjie Wang1,2, Junjun Li5
1Hefei Cancer Hospital of CAS, Institute of Health and Medical Technology, Hefei Institutes of Physical Science, Chinese Academy of Sciences, Hefei, 230031, P. R. China.
Journal of Medical Systems
|May 15, 2025
Summary
Large language models (LLMs) show promise in predicting immunotherapy response for hepatocellular carcinoma (HCC). Gemini-GPT achieved comparable accuracy to senior physicians, offering a potential new tool for clinical decision-making in HCC treatment.
Area of Science:
- Oncology
- Artificial Intelligence
- Medical Imaging Analysis
Background:
- Hepatocellular carcinoma (HCC) is an aggressive malignancy with limited predictive biomarkers for immunotherapy response.
- Large language models (LLMs) offer potential for multimodal data analysis in clinical decision-making.
- The comparative effectiveness of LLMs versus human experts in predicting HCC immunotherapy response is not well-established.
Purpose of the Study:
- To assess the performance of GPT-4, GPT-4o, and Gemini in predicting immunotherapy response in unresectable HCC.
- To compare LLM performance against radiologists and oncologists with varying levels of expertise.
- To evaluate LLM agreement and identify optimal strategies for improved predictive accuracy.
Main Methods:
- Retrospective analysis of 186 patients with unresectable HCC using multimodal data (clinical information and CT images).
- LLMs (GPT-4, GPT-4o, Gemini) were evaluated using zero-shot prompting with 'voting' and 'OR rule' methods.
- Performance metrics included accuracy, sensitivity, area under the curve (AUC), and inter-rater agreement (kappa statistic).
Main Results:
- GPT-4o achieved 65% accuracy and 47% sensitivity using the 'OR rule method', comparable to intermediate physicians.
- Gemini-GPT demonstrated an AUC of 0.69, similar to senior physicians (AUC: 0.72), with 68% accuracy, outperforming junior/intermediate physicians.
- LLMs exhibited higher inter-model agreement (κ = 0.59–0.70) than inter-physician agreement (κ = 0.15 among junior physicians).
Conclusions:
- LLMs, particularly Gemini-GPT, show potential as valuable tools for predicting immunotherapy response in HCC.
- Gemini-GPT's performance is comparable to senior physicians in accuracy, though sensitivity remains lower.
- LLMs demonstrate superior inter-model consistency compared to human expert agreement, suggesting potential for standardized assessment.


