Related Experiment Video
Updated: May 21, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Evaluating the ability of generative AI to standardize TNM staging data in lung cancer
Fatimah Abdulazim Altuhaifa1, Aaron Percy Pereira2, Mustafa Abdulazim Al Tuhaifa3
1Software Development and Research Department, Espresso Raqami, Wollongong, Australia.
Abstract:
This study evaluates the ability of generative artificial intelligence (AI) models to standardize tumor-node-metastasis (TNM) staging data in lung cancer. It assesses how well two large language models, ChatGPT and Gemini, convert information from historical editions of the American Joint Committee on Cancer (AJCC) staging system into the eighth edition. ChatGPT and Gemini were applied to retrospectively standardize TNM staging across multiple AJCC editions using lung cancer data from the Surveillance, Epidemiology, and End Results cancer registry. We used a zero-shot prompting method, employing Python to generate prompts dynamically for each case, and a few-shot variant with examples added to the same inputs. The AI-generated outputs were compared with expert-derived staging to evaluate their accuracy. ChatGPT achieved a micro-accuracy of 52.26% across T, N, M, and stage, slightly outperforming Gemini (50.72%), whereas few-shot prompting improved micro-accuracy to 71.35% and 77.24%, respectively. Both models showed mixed agreement with expert-derived T, N, and M components but often failed to perform the final stage assignment. These results indicate that, although generative AI models may support digital transformation and clinical data standardization, they still face challenges in tasks requiring detailed medical reasoning. Further improvements and validation are needed before relying on generative AI for TNM staging standardization.