A retrospective study using machine learning to develop predictive model to identify rotavirus-associated acute

Sourav Paul1, Minhazur Rahman2, Anutee Dolley3

  • 1Department of Biotechnology, National Institute of Technology, Durgapur, West Bengal, India.

Peerj
|April 18, 2025
PubMed

Insights

Machine learning models can predict rotavirus infection in children using clinical symptoms, aiding diagnosis in resource-limited settings. The Random Forest model showed the highest accuracy at 81.4%.

Area of Science:

  • Pediatric infectious diseases
  • Computational epidemiology
  • Clinical informatics

Background:

  • Rotavirus is a primary cause of severe dehydrating diarrhea in young children globally.
  • Limited access to laboratory diagnostics in hospitals necessitates alternative diagnostic approaches.
  • Machine learning (ML) shows promise for symptom-based disease diagnosis in resource-constrained environments.

Purpose of the Study:

  • To develop an ML predictive model for rotavirus infection using clinical parameters.
  • To avoid reliance on laboratory tests for diagnosis.
  • To support timely and accessible rotavirus diagnosis.

Main Methods:

  • Collected clinical data from 509 children, including symptoms like diarrhea, vomiting, fever, and dehydration.
  • Performed correlation and feature selection (ANOVA F test) to identify important clinical indicators.
  • Trained and compared seven supervised ML models: SVM, KNN, NB, Log_R, RF, DT, and XGBoost.
  • Evaluated model performance using accuracy, precision, recall, specificity, F1, F2, macro F1, and AUC.

Main Results:

  • The Random Forest (RF) model demonstrated superior performance among the seven ML models.
  • RF achieved an accuracy of 81.4%, an F1 score of 86.9%, a macro F1-score of 77.3%, an F2 score of 86.5%, and an AUC of 89%.

Conclusions:

  • ML models can aid in the symptom-based diagnosis of rotavirus-associated acute gastroenteritis in children.
  • These models are particularly valuable in resource-limited settings.
  • Further validation with larger datasets is recommended to optimize sensitivity and specificity for pediatric diarrheal diseases.
Abstract