Related Experiment Video
Updated: Aug 28, 2026

Computer-Aided Three-Dimensional Visualization in the Treatment of Locally Advanced Thyroid Cancer
Published on: June 9, 2023
Assessing the performance of three large language models in thyroid cancer tumour board decision-making
Angus M H White1,2, Kerry A Leyton3, Neil Patel3
1Department of Endocrine and General Surgery, University Hospital of Wales, Heath Park Way, Heath Park, Cardiff, CF14 4XW, UK. angus.white@wales.nhs.uk.
Abstract:
Large language models (LLMs) have potential to support clinical decision-making, but their role in thyroid cancer multidisciplinary team (MDT) meetings remains uncertain. This study evaluated the performance of ChatGPT GPT-5.5, Muse Spark 1.0 (Meta AI) and DeepSeek-V4-Flash in reproducing the management decisions of a regional thyroid cancer MDT. One hundred thyroid cancer cases discussed by a regional MDT were submitted to each LLM using an identical standardised prompt referencing ATA, BTA and UICC guidelines. Recommendations were independently assessed by four consultant endocrine surgeons using a four-point concordance scale (0-3). Consensus scores were determined using the median. Inter-rater agreement was assessed using Fleiss' Kappa. Differences between models were analysed using Friedman and post hoc Wilcoxon signed-rank tests. Three hundred LLM recommendations were assessed. Inter-rater agreement was moderate (Fleiss' Kappa = 0.422, 95% CI 0.369-0.474). ChatGPT achieved the highest proportion of complete concordance with MDT recommendations (69%), followed by Meta AI (60%) and DeepSeek (53%). Clinically acceptable recommendations (scores 2-3) were produced in 93%, 88% and 89% of cases, respectively. Overall concordance differed significantly between models (Friedman χ2(2) = 10.03, p = 0.0066). ChatGPT significantly outperformed Meta AI (p = 0.015) and DeepSeek (p < 0.001), while no difference was observed between Meta AI and DeepSeek (p = 0.384). All three LLMs demonstrated high concordance with consultant-led thyroid cancer MDT decisions. ChatGPT achieved the highest overall concordance, although all models generated clinically acceptable recommendations in most cases. LLMs show promise as adjunctive decision-support tools but require continued clinician oversight and robust governance before routine clinical implementation.
