Multi-centre arousal scoring agreement in the Sleep Revolution
Henna Pitkänen1,2, Sami Nikkonen1,2, Marika Rissanen1,3
1Department of Technical Physics, University of Eastern Finland, Kuopio, Finland.
This study evaluated how consistently different experts identify sleep arousals across multiple centers. Researchers found that manual scoring is often inconsistent, especially for spontaneous events and during wakefulness. These findings suggest that current methods for assessing sleep fragmentation need improvement to ensure reliable clinical and research outcomes.
Area of Science:
- Polysomnography and arousal scoring agreement within clinical sleep medicine
- Neurophysiology and sleep architecture research
Background:
Sleep fragmentation remains a significant challenge in clinical diagnostics due to inconsistent interpretation of physiological signals. Prior research has shown that manual identification of brief interruptions in sleep architecture often lacks standardization. No prior work had resolved the extent of variability across diverse international clinical environments. That uncertainty drove this investigation into how experts align their interpretations of polysomnography data. Standard guidelines exist, yet individual application of these rules frequently diverges among trained professionals. This gap motivated a comprehensive assessment of scoring reliability in a multi-center context. Previous studies often focused on single-site data, limiting the generalizability of their findings. The current analysis addresses these limitations by examining performance across ten expert scorers from seven distinct locations.
Purpose Of The Study:
The aim of this investigation was to quantify the agreement of arousal identification within a multi-center polysomnography framework. Researchers sought to determine if current manual scoring practices provide reliable metrics for clinical assessment. The study addressed the persistent challenge of variability in how experts interpret physiological signals during sleep. By comparing annotations from multiple centers, the team examined the robustness of standardized guidelines. This work was motivated by the need to understand how subjective interpretation impacts the quantification of sleep fragmentation. The authors explored whether specific sleep stages or event types contribute to higher levels of disagreement. They also investigated the frequency of consensus by analyzing clusters of overlapping annotations. This research provides a critical look at the limitations of human-based scoring in modern sleep medicine.
Main Methods:
Review approach involved a multi-center evaluation of expert annotation consistency using 50 full-night recordings. The team recruited ten experienced professionals from seven distinct clinical sites to perform the assessments. Each participant followed the American Academy of Sleep Medicine protocols to identify events. Investigators calculated intraclass correlation coefficients to compare the total frequency of events identified by each expert. They also applied kappa statistics to determine the second-by-second overlap of annotations across the entire duration of the sleep studies. To further characterize the data, the group extracted arousal clusters representing periods where multiple experts identified the same event. This systematic comparison allowed for the quantification of agreement across different sleep stages and event types. The methodology focused on identifying specific areas of uncertainty by comparing individual performance against the collective group consensus.
Main Results:
Key findings from the literature reveal that the overall similarity of arousal indexes was only fair, with an intraclass correlation coefficient of 0.41. Performance between individual scorer pairs ranged widely from 0.04 to 0.88. Respiratory-related events demonstrated better consistency, reaching an intraclass correlation coefficient of 0.65. Spontaneous events showed significantly lower agreement, with an intraclass correlation coefficient of 0.23. The overall second-by-second agreement measured by Fleiss' kappa was 0.40, with individual pair comparisons ranging from 0.07 to 0.68. Agreement improved during deeper sleep stages, reaching 0.53 in N3, while the wake stage showed the lowest reliability at 0.25. Analysis of arousal clusters indicated that over half were identified by only one or two experts. Less than one-third of these clusters received confirmation from at least five of the ten scorers.
Conclusions:
The authors propose that manual identification of sleep interruptions lacks sufficient reliability for consistent clinical application. Synthesis and implications suggest that current assessment methods for sleep fragmentation require substantial refinement. Researchers highlight that spontaneous events and wakefulness periods represent the most challenging scenarios for scoring consensus. The data indicate that variability persists regardless of the standardized guidelines currently in use. This study suggests that relying on individual human interpretation may introduce significant bias into sleep research. The findings imply that automated or more objective criteria are necessary to improve diagnostic accuracy. The authors maintain that the observed low agreement necessitates a re-evaluation of how fragmentation is quantified. Future efforts should prioritize developing more robust metrics to replace existing subjective scoring practices.
Frequently Asked Questions
The researchers report an overall intraclass correlation coefficient of 0.41 for arousal indexes. This indicates only fair agreement across the ten experts involved in the study.
The team utilized Fleiss' kappa to assess second-by-second agreement across the entire recording. They also employed Cohen's kappa to evaluate performance between individual scorer pairs.
The authors report that agreement was notably higher for respiratory-related events, with an intraclass correlation coefficient of 0.65. In contrast, spontaneous arousals showed much lower consistency, reaching only 0.23.
The study utilized 50 full-night polysomnograms for the analysis. These recordings were annotated by ten experts from seven different clinical centers to ensure a diverse dataset.
The researchers observed that Fleiss' kappa values increased from 0.45 in N1 sleep to 0.53 in N3 deep sleep. Conversely, the wake stage exhibited the lowest agreement at 0.25.
The authors propose that manual scoring is generally unreliable for clinical and research purposes. They suggest that systemic changes are necessary to improve the assessment of sleep fragmentation.
Related Concept Videos
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Understanding Sleep
The circadian rhythm, a nearly 24-hour cycle, is deeply influenced by environmental light cues. Light exposure directly affects the hypothalamus, which in turn regulates...
REM Sleep Behavior Disorder
RBD is significantly associated with...
Sleep-Wake Cycles
NREM Sleep
NREM sleep comprises four progressive stages that seamlessly merge:
Optimal Arousal Theory
Inverted U-Shaped Performance Curve
The...
Primary Motives: Sleep, Sex, and Pain Avoidance
Sleep is a fundamental physiological drive that fosters a state of restfulness crucial for several bodily functions. It facilitates body restoration, the process by which the body repairs, rejuvenates, and maintains itself during sleep, including memory...


