Analyzing clinical ratings of performance on pediatric neuropsychological tests
Gregory G Brown1, Jerel E Del Dotto1, John L Fisk1
1a Neuropsychology Division, Department of Psychiatry , Henry Ford Hospital.
Insights
Pediatric neuropsychologists showed fair to excellent agreement in rating child functioning across domains. Agreement was sufficient for combining clinical ratings into broader behavioral measures.
Area of Science:
- Pediatric Neuropsychology
- Developmental Psychology
Background:
- Assessing child neurodevelopmental functioning is complex.
- Standardized rating scales are crucial for consistent evaluation.
Purpose of the Study:
- To evaluate interrater reliability among pediatric neuropsychologists.
- To examine agreement across seven neuropsychological domains for normal and low-birth-weight children.
Main Methods:
- Two studies involving 154 and 41 children, respectively.
- Pediatric neuropsychologists rated children's functioning in specific domains.
- Intraclass correlations (ICCs) were calculated to measure agreement.
Main Results:
- Phase 1 ICCs ranged from .39 (attention) to .85 (intelligence).
- Phase 2 showed similar agreement, with attention agreement increasing to .53.
- Linear modeling indicated differing emphasis on domains by raters.
Conclusions:
- Interrater agreement was generally sufficient for clinical utility.
- Findings support combining individual test scores into global behavioral measures.
- Attention ratings showed the most variability but improved with more raters.
Abstract:
This project examined the agreement among pediatric neuropsychologists when rating the functioning of normal and low-birth-weight children in seven neuropsychological domains. In Phase 1, two neuropsychologists rated 154 children; in Phase 2, three neuropsychologists rated 41 children. Intraclass correlations of agreement in Phase 1 were: attention .39, intelligence .85, auditory/linguistic .82, haptic .70, visual perceptual/visuomotor .78, mnestic .72, and global .61. Intraclass correlations observed in Phase 2 were similar to those found in Phase 1 except that agreement in rating attention increased to .53. Linear modeling of global judgments revealed that two raters emphasized the auditory/linguistic, haptic, and visual perceptual/visuomotor domains in deriving global ratings, whereas the third emphasized the intellectual and attentional domains. In general, agreement among raters was sufficient to justify using clinical ratings to combine individual test scores into more general behavioral measures.


