Feasibility of a German High-Frequency Neonatal and Pediatric Intensive Care Dataset (AIx-Neo-Guard Dataset):
Lena Sophie Olivier1, Camelia Lauterbach Oprea2, André Stollenwerk2
1Division of Neonatology, Department of Pediatric and Adolescent Medicine, Uniklinik RWTH Aachen, Aachen, Germany.
Background:
Available information about existing neonatal and pediatric intensive care datasets is scarce.
Objective:
The objective was to evaluate the feasibility of a monocentric high-frequency dataset of neonatal and pediatric intensive care patients in a cohort study.
Methods:
This study included patients treated in the neonatal and pediatric tertiary care intensive care unit at Uniklinik RWTH Aachen, Germany. The dataset comprised high-frequency data on bedside monitoring and ventilation, CO2 monitoring, laboratory results, important diagnoses and interventions, patient characteristics, and manual annotations (including procedures, eg, intubation and X-rays; and states and illnesses, such as episodes of apnea of prematurity and patient-ventilator asynchronies).
Results:
In total, 400 admissions (227 neonatal and 173 pediatric) from 359 patients were included between September 2024 and January 2026. Of these, 28 neonates were extremely preterm, 34 were very preterm, 68 were moderate to late preterm, and 97 were term infants, with an overall median gestational age of 35 (IQR 31-38) weeks and a median birth weight of 1805 (IQR 1446-3251) g. The pediatric population comprised 34 infants aged younger than 1 year, 45 toddlers aged 1 to 5 years, 43 children aged 6 to 12 years, and 51 adolescents aged 13 to 18 years, with a median age of 10.6 (IQR 1.2-13.0) years and a median admission weight of 22 (IQR 11-50) kg. Overall, 77% (400/517) of all eligible patients were enrolled in the study. The dataset comprised 3600 patient days with more than 20,000 manual annotations (approximately 3 TB) and has been used for multiple use cases.
Conclusions:
The high temporal resolution allows for detailed characterization of patient states. To our knowledge, there are no published descriptions of comparable European high-frequency neonatal or pediatric intensive care datasets. We share information about our dataset to encourage cross-institutional cooperation. The multicenter pooling of data or federated learning increases possible use cases by enabling new research questions or the development of more robust algorithms.


