Related Experiment Video
Updated: Apr 14, 2026

Testing Targeted Therapies in Cancer using Structural DNA Alteration Analysis and Patient-Derived Xenografts
Published on: July 25, 2020
Evaluation of GPT-5, a Large Language Model, in Replicating German Clinical Practice Guideline Recommendations in
Julius Hirsch1,2, Keskanya Subbalekha3, Chatpong Tangmanee4
1Private Oral and Maxillofacial Surgery Practice "Dr. Dr. Hirsch & Söhne", Ulm, Germany.
Background:
Artificial intelligence (AI) technologies, particularly large language models (LLMs) such as ChatGPT, are increasingly utilised in medical education and clinical information retrieval. Nevertheless, their capacity to accurately reproduce recommendations from established clinical practice guidelines (CPGs) has not been thoroughly examined. The present study evaluated the concordance between responses generated by GPT-5 and recommendations contained in German CPGs addressing oral potentially malignant disorders (OPMDs) and oral carcinomas (OCs).
Methods:
A cross-sectional analytical comparison was performed between GPT-5 outputs and German CPG recommendations available as of October 2025. Individual guideline statements were entered verbatim into GPT-5, which was asked to confirm or reject the statements. To assess methodological robustness, inverted versions of the same statements were additionally tested. GPT-5 was accessed through the free version without internet connectivity to ensure that responses originated solely from the model's internal training data. Accuracy was defined as the proportion of correctly classified statements. Concordance between guideline content and model responses was quantified using Cohen's 𝝹.
Results:
Two German CPGs comprising 111 recommendations were included: the S2k guideline for OPMDs (15 recommendations) and the S3 guideline for OCs (96 recommendations). GPT-5 correctly affirmed all authentic recommendations and rejected all inverted statements. Agreement between guideline statements and GPT-5 responses was perfect when the original recommendations were analysed (𝝹 = 1.0) and remained very high when both original and inverted statements were evaluated jointly (𝝹 = 0.96). The majority of references cited within the guidelines were published in English (> 93%) and originated from outside Germany (> 77%).
Conclusion:
When guideline recommendations were presented verbatim, GPT-5 demonstrated complete concordance with German oral oncology CPGs. These findings indicate that the model is capable of recognising and retrieving established guideline information. However, this experimental design evaluates recognition of existing statements rather than autonomous clinical reasoning. At present, LLMs should therefore be regarded primarily as educational and informational tools rather than a replacement for expert clinical judgement in oral oncology.
Related Concept Videos
Cancer Therapies
However, cancer treatments can pose several challenges, as therapies used to kill cancer cells are generally also toxic to normal cells. Moreover, cancer cells mutate rapidly and can develop resistance to chemical agents or radiation therapy. Besides, all types of cancer cells may not respond to the same therapy. Some cancer cells respond to one...
Treatment Resistant Cancers
Treatment Resistent Cancers