Interpretable multi-center machine learning model driven by facial image features for non-invasive early risk
Yulin Shi1,2, Yi Chun2, Shuyi Zhang2
1Teaching Experiment and Training Center, Academic Affairs Office, Shanghai University of Traditional Chinese Medicine, Shanghai, China.
Background:
Early identification of lung cancer is critical for improving patient survival and prognosis. Conventional pathological biopsy is invasive, and computed tomography (CT) involves ionizing radiation risks. Facial imaging, as a non-invasive and accessible biological signal, is a promising novel approach for lung cancer auxiliary diagnosis.
Objective:
To develop and validate a non-invasive lung cancer auxiliary diagnosis model based on facial image features, and explore the value of facial imaging in lung cancer early screening.
Methods:
Multi-center patients with benign pulmonary nodules and lung cancer were enrolled in this study. Facial images were collected via the TFDA-1 Digital Tongue and Face Diagnosis Instrument, with 124 features extracted. Logistic regression was used for feature selection and simple correlation analysis of facial image features, based on which four machine learning models (XGBoost, LightGBM, SVM, GBDT) were constructed. Model training and optimization were performed using 10-fold cross-validation combined with grid search for hyperparameter tuning. Model performance was comprehensively evaluated using Accuracy, Precision, Sensitivity, Specificity, F1-Score, area under the curve (AUC), average precision (AP), and Brier Score; pairwise comparisons of AUCs among models were conducted using the DeLong test. Clinical utility was assessed using calibration curves and decision curves; model interpretability was analyzed via the SHapley Additive exPlanations (SHAP) method; and generalization capability was validated using an independent external validation cohort.
Results:
The XGBoost model achieved the best overall performance, with an AUC of 0.900 and accuracy of 0.807 in the internal test set, and an AUC of 0.906 and accuracy of 0.813 in the external validation set, demonstrating favorable generalization stability. DeLong test showed that XGBoost achieved the highest AUC in the internal test set without significant differences among models (all P ≥ 0.05). In external validation, its AUC was significantly higher than GBDT, LightGBM (P < 0.001) and SVM (P < 0.05).
Conclusion:
We successfully constructed a non-invasive auxiliary screening model for lung cancer using facial image features. Facial imaging shows significant value in early lung cancer screening, providing a novel, accessible strategy to improve the popularization and availability of lung cancer early screening in clinical practice.
More Related Videos
07:53Three-Dimensional Reconstruction for the Whole Lung with Early Multiple Pulmonary Nodules
Published on: October 13, 2023
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
