A Machine Learning Based Framework to Identify and Classify Non-alcoholic Fatty Liver Disease in a Large-Scale
Weidong Ji1, Mingyue Xue2, Yushan Zhang3
1Department of Medical Information, Zhongshan School of Medicine, Sun Yat-sen University, Guangzhou, China.
Frontiers in Public Health
|April 21, 2022
Summary
Machine learning models accurately screen for non-alcoholic fatty liver disease (NAFLD) using physical exam data. XGBoost identified key factors like BMI and age for early NAFLD detection and management.
Area of Science:
- Hepatology
- Medical Informatics
- Public Health
Background:
- Non-alcoholic fatty liver disease (NAFLD) is a prevalent global health concern with limited effective treatments.
- Accurate and efficient screening methods are crucial for managing the growing burden of NAFLD.
- Machine learning (ML) offers potential for developing advanced diagnostic tools in healthcare.
Purpose of the Study:
- To develop and validate machine learning models for the accurate screening of non-alcoholic fatty liver disease (NAFLD).
- To identify key predictive factors for NAFLD using a large-scale dataset.
- To assess the performance of different ML algorithms in NAFLD screening.
Main Methods:
- Utilized data from 304,145 adults undergoing national physical examinations, including questionnaire and physical measurement parameters.
- Applied Absolute Shrinkage and Selection Operator (LASSO) for feature selection from candidate covariates.
- Developed and compared four ML algorithms (XGBoost, etc.) for NAFLD screening, selecting the best-performing classifier.
Main Results:
- XGBoost demonstrated superior performance with an accuracy of 0.880, precision of 0.801, recall of 0.894, F-1 score of 0.882, and AUC of 0.951.
- Key covariates identified for NAFLD screening importance were BMI, age, waist circumference, gender, type 2 diabetes, gallbladder disease, smoking, hypertension, dietary status, physical activity, and dietary preferences.
- The developed ML models facilitate early identification and classification of NAFLD.
Conclusions:
- Machine learning classifiers, particularly XGBoost, provide an effective tool for the early identification and classification of NAFLD.
- The identified covariate importance ranking offers valuable insights for NAFLD prevention and treatment strategies.
- These ML models are particularly beneficial for resource-limited settings to improve NAFLD screening accessibility.
Keywords:
LASSOmachine learningnon-alcoholic fatty liver disease (NAFLD)predictive modelsscreening modelMore Related Videos
07:03Author Spotlight: Establishing MASLD Cell Models for Investigating Disease Mechanisms and the Lipid-Lowering Effects of Koumiss
Published on: July 19, 2024
1.1K
08:41Novel In Vivo Micro-Computed Tomography Imaging Techniques for Assessing the Progression of Non-Alcoholic Fatty Liver Disease
Published on: March 24, 2023
1.3K
