Related Experiment Video
Updated: Jun 8, 2026

Automated Dissection Protocol for Tumor Enrichment in Low Tumor Content Tissues
Published on: March 29, 2021
Automated Tumor International Classification of Diseases Coding of Real-World Pathology Reports Using Self-Hosted
Kamyar Arzideh1,2, René Hosch2,3, Amin Turki4,5,6
1Central IT Department, Data Integration Center, University Hospital Essen, Essen, Germany.
Purpose:
Manual coding of pathology reports with International Classification of Diseases for Oncology (ICD-O)-3 codes is time-consuming, error-prone, and resource-intensive for health care institutions. To evaluate the performance of multiple state-of-the-art large language models (LLMs) in extracting ICD-O-3 topography and morphology codes from real-world pathology reports and assess their potential for clinical implementation, this study compares the performance of state-of-the-art open-source models in multiple evaluation setups.
Methods:
We analyzed 21,364 pathology reports from 10,823 patients documented between 2013 and 2025 at a large German hospital. Five LLMs were evaluated: Llama-3.3-70B-Instruct, DeepSeek-R1-Distill-Llama (8B and 70B variants), Qwen3-235B-A22B, and Gemma-3-12B-it. All models were deployed on secured private information technology hospital infrastructure. Three different prompts were developed for topography extraction (with and without anatomic context) and morphology extraction. Performance was evaluated using exact code matches and first three-position matches.
Results:
For exact ICD-O topography code prediction, Qwen3-235B-A22B achieved the highest performance (microaverage F1: 71.6%), whereas Llama-3.3-70B-Instruct performed best at predicting the first three characters (micro-average F1: 84.6%). For morphology codes, DeepSeek-R1-Distill-Llama-70B outperformed other models (exact microaverage F1: 34.7%; first three characters' microaverage F1: 77.8%). Large disparities between micro- and macroaverage F1-scores indicated poor generalization to rare conditions.
Conclusion:
Although LLMs demonstrate promising capabilities as support systems for expert-guided pathology coding, their performance is not yet sufficient for fully automated, unsupervised use in routine clinical workflows. LLMs showed poor performance on rare conditions, heavy dependence on contextual information, and substantially lower scores for morphology versus topography classification.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Related Concept Videos
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll