Prevalence of Fatty Liver among Children under Multiple Machine Learning Models

Yunlong Lu1, Wenyu Li1, Xiangbo Gong1

  • 1From the School of Mathematics and Statistics, Beihua University, Jilin, China, and the Departments of Mathematics and Physics and Biology and Chemistry, Texas A&M International University, Laredo, Texas.

Insights

Machine learning models identified key factors for childhood fatty liver disease in South Texas. Higher body mass index in children is strongly associated with an increased probability of developing this condition.

Area of Science:

  • Pediatric Gastroenterology
  • Medical Informatics
  • Biostatistics

Background:

  • Childhood fatty liver disease is a growing concern.
  • Ultrasound data analysis is crucial for diagnosis.
  • Machine learning offers novel approaches for disease prediction.

Purpose of the Study:

  • To identify factors contributing to fatty liver in children using ultrasound data.
  • To develop machine learning models for predicting childhood fatty liver.
  • To inform prevention and treatment strategies for pediatric fatty liver disease.

Main Methods:

  • Utilized the CatBoost algorithm for feature selection.
  • Employed grid search for parameter optimization.
  • Developed and evaluated binary classification models (logistic regression, CatBoost) for fatty liver prediction in obese children.
  • Compared model performance using AUC, precision, accuracy, recall, and F1 score.

Main Results:

  • Selected features included body mass index, height, liver size, kidney volume, glomerular filtration rate, and liver diameter.
  • Higher body mass index correlated with increased fatty liver probability in children.
  • Machine learning models demonstrated predictive capabilities for fatty liver in obese children.

Conclusions:

  • Logistic regression and CatBoost models predict a high probability of fatty liver in severely obese children (74.47%-92.22%) and obese children (73.45%-85.41%).
  • Boys showed a slightly higher mean probability of fatty liver compared to girls (3.00%-3.95% difference).
  • Machine learning models provide valuable insights into childhood fatty liver risk factors and prevalence.
Abstract