Related Experiment Video
Updated: Jul 18, 2026

Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
Can ChatGPT and Gemini justify brain CT referrals? A comparative study with human experts and a custom prediction
Jaka Potočnik1, Edel Thomas2, Dearbhla Kearney3
1University College Dublin School of Medicine, Dublin, Ireland. jaka.potocnik@ucd.ie.
Background:
The poor uptake of imaging referral guidelines in Europe results in a substantial amount of inappropriate computed tomography (CT) scans. Publicly available chatbots, ChatGPT and Gemini, offer an alternative for justifying real-world referrals. Recent research reports high ChatGPT accuracy when analysing American College of Radiology Appropriateness Criteria variants. We compared the chatbots' performance in interpreting, justifying, and suggesting alternative imaging for unstructured adult brain CT referrals in accordance with the European Society of Radiology iGuide. Our prediction model for automated iGuide categorisation of referrals was also compared against the chatbots.
Methods:
The iGuide justification of 143 real-world CT brain referrals, used to evaluate a prediction model, was analysed by two radiographers and radiologists. ChatGPT-4's and Gemini's imaging recommendations and pathology suspicions were compared with those of humans, with respect to referral completeness. Inter-rater reliability with κ statistics determined the agreement between entities.
Results:
Chatbots' performance was limited (κ = 0.3) but improved for more complete referrals. The prediction model outperformed the chatbots in justification analysis (κ = 0.853). The chatbots' interpretations of complete referrals were highly consistent (49/52, 94.2%). The agreement regarding alternative imaging was high for both complete and ambiguous referrals, with ChatGPT and Gemini correctly identifying imaging modality and anatomical region in 83/96 (86.5%) and 81/96 (84.4%) cases, respectively.
Conclusion:
The chatbots' ability to analyse the justification of adult brain CT referrals is limited to complete referrals, unlike our prediction model. Further research is needed to confirm these findings for other types of CT scans and modalities.
Relevance Statement:
ChatGPT and Gemini exhibit potential in justifying free text brain CT referrals; however, further improvements are required to handle real-world referrals of varying quality.
Key Points:
Custom prediction model's justification analysis strongly aligns with iGuide and surpasses chatbots. Chatbots incorrectly justified almost one-half of all CT brain referrals. Chatbots have limited performance in justifying ambiguous CT brain referrals. Chatbot performance improved when referrals were detailed and included suspected pathology.
More Related Videos
09:03Radiotracer Administration for High Temporal Resolution Positron Emission Tomography of the Human Brain: Application to FDG-fPET
Published on: October 22, 2019
07:57Positron Emission Tomography-based Dose Painting Radiation Therapy in a Glioblastoma Rat Model using the Small Animal Radiation Research Platform
Published on: March 24, 2022