Investigation of predictive factors for fatty liver in children and adolescents using artificial intelligence

Aliakbar Sayyari1, Amin Magsudy2, Yasamin Moeinipour3

  • 1Pediatric Gastroenterology, Hepatology and Nutrition Research Center, Research Institute for Children's Health, Shahid Beheshti University of Medical Sciences, Tehran, Iran.

Frontiers in Pediatrics
|August 27, 2025
PubMed

Insights

Machine learning accurately predicts childhood non-alcoholic fatty liver disease (NAFLD). The CatBoost model showed high accuracy, aiding early diagnosis and improving outcomes for pediatric NAFLD.

Area of Science:

  • Pediatric Hepatology
  • Medical Informatics
  • Machine Learning in Medicine

Background:

  • Childhood obesity is a global health concern, increasing the incidence of non-alcoholic fatty liver disease (NAFLD), the most prevalent liver condition in children.
  • Liver biopsy, the current standard for NAFLD diagnosis, can be invasive. Early detection through advanced methods is crucial for better patient prognosis.
  • Machine learning (ML) offers a promising avenue for developing non-invasive, early diagnostic tools for pediatric NAFLD.

Purpose of the Study:

  • To identify key predictive factors for NAFLD in pediatric populations using ML models.
  • To evaluate the efficacy of various ML algorithms in diagnosing NAFLD based on liver biopsy outcomes.
  • To assess predictive performance for specific NAFLD histological features like fibrosis, steatosis, and ballooning.

Main Methods:

  • Analysis of data from 659 children with suspected NAFLD who underwent liver biopsy between 2011 and 2023.
  • Data preprocessing involved one-hot encoding for categorical variables and standardization for numerical features.
  • Training and evaluation of ML models including CatBoost, AdaBoost, Random Forest, and GradientBoosting using cross-validation and metrics like accuracy, precision, recall, F1 score, and ROC-AUC.

Main Results:

  • The CatBoost Classifier achieved the highest predictive accuracy (91.8%) and ROC-AUC (92.3%) in cross-validation for NAFLD diagnosis.
  • Adjusted models demonstrated improved performance, with CatBoost's F1 score increasing from 83% to 89% (AUC: 0.86-0.92).
  • Other models like GradientBoosting and Bernoulli Naive Bayes also showed enhanced predictive capabilities after adjustments.

Conclusions:

  • Machine learning models, especially CatBoost, exhibit significant potential for accurate and early diagnosis of NAFLD in children.
  • These findings suggest ML can serve as a valuable tool to support clinical decision-making in pediatric NAFLD.
  • Further development and validation of ML algorithms could lead to improved diagnostic strategies and patient management for childhood NAFLD.
Abstract