Related Experiment Video
Updated: Aug 5, 2026

Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
Inter-rater reliability and agreement of the Bayley-4 in a multidisciplinary team
Sarah E Hall1,2,3, Samudragupta Bora4,5, Caroline Alexander6,7
1Curtin University, Perth, WA, Australia. sarah.hall@curtin.edu.au.
Insights
The Bayley-4 assessment shows excellent reliability when used by multidisciplinary teams. Structured training ensures consistent results in early childhood development evaluations.
Area of Science:
- Neurodevelopmental assessment
- Pediatric allied health
- Psychometrics
Background:
- The Bayley Scales of Infant and Toddler Development are crucial for early development assessment.
- Limited evidence exists on inter-rater reliability in multidisciplinary settings for the Bayley-4.
- This study addresses the need for reliability data in diverse clinical teams.
Purpose of the Study:
- To evaluate the inter-rater reliability and agreement of the Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-4).
- To assess reliability when administered by a multidisciplinary allied health team at age two years.
- To provide independent evidence of Bayley-4 reliability beyond its standardization sample.
Main Methods:
- 100 children aged two years from the Early Moves study in Perth, Australia, were assessed.
- 18 trained clinicians from various allied health disciplines independently scored assessments in real time.
- Inter-rater reliability was analyzed using intraclass correlation coefficients (ICCs) and Bland-Altman analysis.
Main Results:
- Excellent inter-rater reliability was found across Bayley-4 Cognitive, Language, and Motor composite scores and subtest scaled scores (ICC range: 0.96-1.00).
- Mean differences between raters were negligible, with narrow limits of agreement (<±0.5 SD).
- High scoring consistency was achieved by the multidisciplinary team, including those new to the Bayley scales.
Conclusions:
- The Bayley-4 demonstrates excellent inter-rater reliability and clinically acceptable agreement in a multidisciplinary allied health team context.
- Formal training and standardization procedures are key to achieving high reliability in neurodevelopmental assessments.
- Findings support the use of multidisciplinary teams for Bayley-4 assessments and emphasize the importance of ongoing training and feedback.
Abstract:
The Bayley Scales of Infant and Toddler Development are widely used to assess early development, yet evidence for inter-rater reliability in multidisciplinary contexts remains limited. This study evaluated inter-rater reliability and agreement of the newest edition, Bayley-4, at age two years when administered by a multidisciplinary allied health team within a longitudinal cohort. Participants were 100 children comprising a randomly selected 5% subsample of the population-based Early Moves study in Perth, Australia. Assessments were independently double scored in real time by 18 trained clinicians (seven physiotherapists, six occupational therapists, three speech pathologists and two psychologists). Inter-rater reliability was evaluated using intraclass correlation coefficients (ICCs) with 95% confidence intervals, and agreement using Bland-Altman analysis. Inter-rater reliability was excellent across Cognitive, Language and Motor composite scores and subtest scaled scores (ICC range: 0.96-1.00). Mean differences between raters were negligible, and limits of agreement were narrow and within predefined clinically acceptable thresholds (<±0.5 SD). These findings demonstrate that excellent inter-rater reliability and agreement for the Bayley-4 can be achieved within a large multidisciplinary allied health team when supported by formal structured training and ongoing standardisation procedures. Further research should evaluate reliability in clinical populations and routine service contexts. IMPACT: The Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-4) demonstrated excellent inter-rater reliability and clinically acceptable agreement at age two years. High scoring consistency was achieved across a large multidisciplinary allied health team. Reliability was maintained even among clinicians without prior Bayley experience following structured training and standardised protocols. This study provides independent evidence of Bayley-4 inter-rater reliability beyond the standardisation sample. Findings support multidisciplinary assessment models and highlight the importance of accredited training, supervised practice and ongoing feedback processes for maintaining reliable neurodevelopmental assessment outcomes.