Related Experiment Videos
Machine Learning on the Pediatric Intensive Care (PIC) Database: PIC-Powered Prediction
Hammad Ashraf Ganatra1, Daniah Shamim2, Shawn B Sood3
1Pediatric Critical Care Medicine, Cleveland Clinic Children's Hospital, Cleveland, OH 44195, USA.
Abstract:
Machine learning (ML) applied to pediatric intensive care data could enable earlier risk stratification and more precise decision support, but limited shareable pediatric datasets have constrained progress. The Pediatric Intensive Care database (PIC) on PhysioNet provides a de-identified, bilingual electronic health record resource spanning 2010-2018. In this narrative review we synthesize ten PIC-based ML studies identified through forward citation tracking on the original PIC publication and PhysioNet dataset record in PubMed, Web of Science, and Google Scholar. For each study we abstracted the clinical question, cohort, label construction, feature engineering, model class, validation design, calibration reporting, and release of reproducibility artifacts. The reviewed studies addressed catheter-associated thrombosis, sepsis, in-hospital mortality, and organ-dysfunction phenotyping. We organize the synthesis around four substrate properties of PIC: irregular time series, constructed labels, single-center temporal drift, and bilingual identifier semantics. The review advances three claims. First, upstream design choices appear to drive performance at least as much as classifier selection across the studies reviewed. Second, within-site discrimination metrics are insufficient without calibration, decision-curve analysis, and deployment-realistic validation. Third, PIC should be viewed as the seed for a collaborative pediatric ICU data ecosystem.