Related Experiment Video
Updated: May 11, 2026

Basics of Multivariate Analysis in Neuroimaging Data
Published on: July 24, 2010
A multivariable model for improving the identification of cerebral palsy cases in administrative health data
Peter M Socha1, Maryam Oskoui2, Jennifer A Hutcheon3
1Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, Canada.
Insights
This study developed a model to better identify cerebral palsy cases in health records. The new method showed improved accuracy compared to existing algorithms, though some misclassification persists.
Area of Science:
- Medical Informatics
- Public Health
- Pediatric Neurology
Background:
- Accurate identification of cerebral palsy (CP) is crucial for timely intervention and resource allocation.
- Administrative health data offer a vast resource but often contain misclassified cases.
- Existing algorithms for CP case identification in administrative data have limitations.
Purpose of the Study:
- To enhance the accuracy of identifying cerebral palsy (CP) cases within population-based administrative health datasets.
- To develop and validate a predictive model for CP detection using logistic regression and ICD codes.
- To compare the performance of the new model against established CP identification algorithms.
Main Methods:
- Utilized a population-based cerebral palsy registry in Quebec, Canada, including children born 1999-2002.
- Analyzed hospitalization and physician billing records up to 2012 for children with and without CP.
- Employed logistic regression modeling with International Classification of Diseases (ICD) codes for related conditions, assessing performance via ROC and PR curves.
Main Results:
- The developed model achieved an area under the ROC curve of 0.98 and PR curve of 0.73.
- At comparable specificity levels, the model demonstrated 1-14 percentage-points higher sensitivity than existing algorithms.
- Higher sensitivity was observed with longer follow-up, combined data sources, and for preterm infants.
Conclusions:
- The novel model significantly improved the identification of cerebral palsy cases in administrative health data.
- Despite improvements, residual misclassification of cerebral palsy cases remains a challenge.
- Findings suggest potential for optimized CP case ascertainment in large-scale health databases.
Purpose:
To improve the identification of cerebral palsy cases in administrative health data.
Methods:
We included all children in a population-based cerebral palsy registry in Quebec, Canada, born from 1999 through 2002, and a sample of children without cerebral palsy. Population-based hospitalization and physician billing records through 2012 were obtained for all children. We used logistic regression to model the probability of cerebral palsy, using International Classification of Diseases codes for related diseases. We reported receiver operating characteristic (ROC) and precision-recall (PR) curves, and compared the accuracy to that of existing algorithms. We also reported the accuracy of cerebral palsy codes by age, data source, and gestational age at birth.
Results:
The area under the ROC and PR curves of our model were 0.98 (95 % CI: 0.97-0.99) and 0.73 (95 % CI: 0.63-0.79), respectively. Cut-offs with a similar specificity to existing algorithms yielded sensitivities that were 1-14 %age-points higher. The sensitivity of cerebral palsy codes was higher (and the specificity was lower) with longer follow-up times since birth, when using both hospitalization and billing records, and among children born preterm.
Conclusions:
Our model improved identification of cerebral palsy cases in administrative data, but residual misclassification remained.

