Related Experiment Videos
Data for Machine Learning Models: Clinical Experience Determines Interrater Reliability in Waveform Anomaly Detection
Josef Škola1,2, Zbyšek Posel3, Petr Waldauf4,5
1Department of Anaesthesiology, Intensive Care and Emergency Medicine, Bulovka University Hospital and Third Medical Faculty, Charles University, Budinova 67/2, 180 00, Prague 8, Czechia. josef.skola@bulovka.cz.
Background:
High-frequency intracranial pressure (ICP) and arterial blood pressure (ABP) waveforms are increasingly utilized for advanced physiological monitoring and as inputs for machine learning (ML) models. The accuracy of derived parameters relies on input data quality, and anomaly contaminated signals may impair model performance. Reliable anomaly annotation is therefore essential. Despite progress in automated labeling, human annotations remain the "ground truth" for ML training approaches. Human labeling is subjective and may vary depending on the signal type and the annotator's experience. This study aimed to quantify the interrater reliability of anomaly annotations in ICP and ABP waveforms.
Methods:
Anonymized, synchronized ICP and ABP waveforms were analyzed retrospectively. From a single-center neurocritical care database, 100 h of recordings (32,760 10 s segments) were selected using predefined criteria. Six annotators (three neurocritical care consultants, two residents, one medical student) annotated anomalies under identical conditions. Interrater reliability was assessed for ICP and ABP using multiple metrics [Gwet's AC2, Fleiss' κ, prevalence- and bias-adjusted kappa (PABAK), positive/negative agreement, and intraclass correlation coefficient, ICC(2,k)], with segment-level and patient-level cluster-robust 95% confidence intervals. A broader anomaly definition, rather than artifact-based labeling, was used to capture variability in judgments of signal fidelity.
Results:
Anomaly prevalence was low (3.99% for ICP, 5.83% for ABP). Overall agreement was high for both signals (Gwet's AC2 ≥ 0.98). However, anomaly specific agreement differed by modality and experience. For ICP, senior annotators achieved higher positive agreement than mixed-experience teams (+ 12.8%), whereas ABP annotation showed no significant experience effect. Across all annotators, ABP exhibited higher anomaly specific concordance, with similar overall agreement.
Conclusions:
Robust interrater reliability in waveform anomaly annotation is achievable in routine neurocritical care; however, reliability remained signal- and experience-dependent. ICP annotation benefits from expert training and supervision, while ABP can be reliably annotated by mixed-experience teams. These findings support signal-specific annotation strategies and provide an empirical basis for developing ML-based anomaly detection systems.