Related Experiment Video
Updated: Aug 14, 2026

Electrophysiological Measurements and Analysis of Nociception in Human Infants
Published on: December 20, 2011
Infant polysomnography: reliability. Collaborative Home Infant Monitoring Evaluation (CHIME) Steering Committee
D H Crowell1, L J Brooks, T Colton
1Kapiolani Medical Center for Women and Children, Honolulu, Hawaii 96826, USA.
Insights
Infant polysomnography (IPSG) scoring reliability was enhanced through detailed training and modified criteria, achieving substantial agreement for sleep parameters and states. This ensures IPSG data is a dependable source for clinical and research applications.
Area of Science:
- Pediatric Sleep Medicine
- Biomedical Data Analysis
- Clinical Trial Methodology
Background:
- Infant polysomnography (IPSG) is crucial for diagnosing infant sleep and breathing disorders.
- Subjectivity in IPSG data analysis necessitates reliable scoring by clinicians.
- Establishing inter-rater and intra-rater reliability is vital for data integrity.
Purpose of the Study:
- To test the hypothesis that infant sleep parameters (SP) and sleep states (SS) can be reliably scored with substantial agreement (kappa >= 0.61).
- To develop and implement a reliability training and evaluation process for IPSG scoring within the CHIME study.
- To assess the reliability of experienced investigators and trainees in scoring IPSG data.
Main Methods:
- Utilized the kappa statistic to evaluate inter- and intra-rater reliability for SP and SS.
- Scored 408 epochs of IPSG data from diverse infant populations (healthy, preterm, apnea, SIDS siblings).
- Implemented iterative refinement of scoring criteria and training based on initial reliability assessments.
Main Results:
- Initial inter-rater reliability for SS showed moderate agreement (kappa = 0.45-0.58).
- Following criteria modification and retraining, subsequent analysis demonstrated substantial agreement.
- Final kappa values for SS (0.68) and SP (0.62-0.76) indicated high reliability among trained scorers.
Conclusions:
- The hypothesis that infant SP and SS can be reliably scored was supported.
- IPSG is a reliable source for clinical and research data when supported by robust reliability metrics.
- Strict scoring guidelines and comprehensive training are essential for maximizing IPSG reliability.
Abstract:
Infant polysomnography (IPSG) is an increasingly important procedure for studying infants with sleep and breathing disorders. Since analyses of these IPSG data are subjective, an equally important issue is the reliability or strength of agreement among scorers (especially among experienced clinicians) of sleep parameters (SP) and sleep states (SS). One basic issue of this problem was examined by proposing and testing the hypothesis that infant SP and SS ratings can be reliably scored at substantial levels of agreement, that is, kappa (kappa) > or = 0.61. In light of the importance of IPSG reliability in the collaborative home infant monitoring evaluation (CHIME) study, a reliability training and evaluation process was developed and implemented. The bases for training on SP and SS scoring were CHIME criteria that were modifications and supplements to Anders, Emde, and Parmelee (10). The kappa statistic was adopted as the method for evaluating reliability between and among scorers. Scorers were three experienced investigators and four trainees. Inter- and intrarater reliabilities for SP codes and SSs were calculated for 408 randomly selected 30-second epochs of nocturnal IPSG recorded at five CHIME clinical sites from healthy full term (n = 5), preterm (n = 4), apnea of infancy (n = 2), and siblings of the sudden infant death syndrome (SIDS) (n = 4) enrolled subjects. Infant PSG data set 1 was scored by both experienced investigators and trained scorers and was used to assess initial interrater reliability. Infant PSG data set 2 was scored twice by the trained scorers and was used to reassess inter-rater reliability and to assess intrarater reliability. The kappa s for SS ranged from 0.45 to 0.58 for data set 1 and represented a moderate level of agreement. Therefore, rater disagreements were reviewed, and the scoring criteria were modified to clarify ambiguities. The kappa s and confidence intervals (CIs) computed for data set 2 yielded substantial inter-rater and intrarater agreements for the four trained scorers; for SS, the kappa = 0.68 and for SP the kappa s ranged from 0.62 to 0.76. Acceptance of the hypothesis supports the conclusion that the IPSG is a reliable source of clinical and research data when supported by significant kappa s and CIs. Reliability can be maximized with strictly detailed scoring guidelines and training.

