Related Experiment Video
Updated: Jun 12, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Diagnostic Performance of Contemporary Large Language Models on Free-Text Histopathologic Descriptions in Oral and
Sohaib Shujaat1, Arun Gopinathan Pillai2, Marryam Riaz3
1King Abdullah International Medical Research Center, Department of Maxillofacial Surgery & Diagnostic Sciences, College of Dentistry, King Saud Bin Abdulaziz University for Health Sciences, Ministry of National Guard Health Affairs, Riyadh, Kingdom of Saudi Arabia. sohaib.shujaat941@gmail.com.
Large language models (LLMs) show moderate accuracy in diagnosing oral and maxillofacial pathology (OMFP) from text narratives. Performance is better for distinctive microscopic features than for correlation-dependent entities.
Area of Science:
- Oral and Maxillofacial Pathology
- Artificial Intelligence in Medicine
- Natural Language Processing
Background:
- Histopathologic narrative descriptions are crucial for OMFP diagnosis.
- Evaluating the diagnostic capabilities of Large Language Models (LLMs) in specialized medical fields is an emerging area of research.
- Benchmarking LLM performance requires standardized methodologies and diverse case sets.
Purpose of the Study:
- To benchmark the diagnostic accuracy of contemporary LLMs using text-only histopathologic descriptions in OMFP.
- To assess inter-model agreement among different LLMs for OMFP diagnosis.
- To evaluate how LLM performance varies based on diagnostic dependency and category.
Main Methods:
- A retrospective study included 155 de-identified OMFP cases with edited histopathologic narratives and reference diagnoses.
- Three general-purpose LLMs (ChatGPT-5.0, Gemini-2.5-Pro, Claude-Opus-4.1) were queried using a zero-shot prompt for a single definitive diagnosis.
- Case-level accuracy was the primary outcome, with paired comparisons using McNemar's test and inter-model agreement assessed with Cohen's κ.
Main Results:
- Overall accuracies were 83.2% (ChatGPT-5.0), 77.4% (Gemini-2.5-Pro), and 72.3% (Claude-Opus-4.1).
- All models performed better on histology-sufficient (HS) lesions than correlation-dependent (CD) lesions.
- Highest accuracy was observed in malignant, hematolymphoid, and immune-mediated OMFP categories.
Conclusions:
- Contemporary LLMs demonstrate moderate diagnostic accuracy for OMFP histopathologic narratives, especially for lesions with distinct microscopic features.
- LLM performance declines for CD entities, highlighting the continued need for clinicoradiologic context.
- LLMs are not yet suitable as standalone diagnosticians but show potential for supportive roles in OMFP education and informatics under supervision.