Related Experiment Video
Updated: Jul 16, 2026

07:13
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Machine Learning-Based Prediction of Masaoka-Koga Stage and WHO Histological Risk Group in Thymic Epithelial Tumors
Konstantinos Kitrou1, Georgios Mandrakis2, Georgios Tsirogiannis3
1Department of Business Administration, University of Patras, 26504 Rio, Achaia, Greece.
Diagnostics (Basel, Switzerland)
|July 15, 2026
Summary
Machine learning models using immunohistochemical H-scores accurately predict thymic epithelial tumor (TET) stages and risk groups. This approach aids in classifying these mediastinal neoplasms, potentially reducing diagnostic variability.
Area of Science:
- Oncology
- Pathology
- Computational Biology
Background:
- Thymic epithelial tumors (TETs) are primary anterior mediastinal neoplasms.
- TET classification involves anatomical staging (Masaoka-Koga) and risk stratification (WHO), relying on expert pathology and facing interobserver variability.
- This study explores machine learning for objective TET classification using quantitative immunohistochemistry (IHC).
Purpose of the Study:
- To apply supervised machine learning (ML) to quantitative IHC H-score profiles for predicting Masaoka-Koga stage and World Health Organization (WHO) risk group in TETs.
- To identify robust IHC-based signatures for TET classification.
- To assess the potential of ML in reducing diagnostic subjectivity in TET pathology.
Main Methods:
- Logistic regression (LR) and XGBoost were employed on 19 biomarkers' H-score profiles.
- Masaoka-Koga stage prediction involved 81 patients with SMOTE oversampling.
- WHO risk group prediction included 89 patients without oversampling; cross-endpoint analysis was also performed.
Main Results:
- Logistic regression outperformed XGBoost for both classification tasks.
- An optimal Masaoka-Koga model (EphA6, YAP, HDAC4 H-scores) achieved an AUC of 0.756.
- An optimal WHO model (TAZ, EphA6, YAP H-scores) reached an AUC of 0.936, with the Masaoka-Koga triad predicting WHO risk with an AUC of 0.901.
Conclusions:
- Quantitative IHC H-score profiling combined with supervised ML can identify biologically interpretable signatures for TET classification.
- The developed models show high predictive performance for both staging and risk stratification.
- Prospective external validation is necessary before clinical implementation of these ML-based approaches.