Related Experiment Video
Updated: Sep 21, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Novel Pediatric Height Outlier Detection Methodology for Electronic Health Records via Machine Learning With
Rodney A Sparapani1, Bi Q Teng1, Julia Hilbrands2
1From the Division of Biostatistics, Medical College of Wisconsin, Milwaukee, WI.
Insights
A new automated method reliably identifies height outliers in children's electronic health records (EHR). This approach ensures accurate body proportionality measurements, improving pediatric health assessments.
Area of Science:
- Pediatric Endocrinology
- Biostatistics
- Health Informatics
Background:
- Accurate height measurements are crucial for monitoring child growth and development.
- Identifying outliers in electronic health records (EHR) is essential for data integrity.
Purpose of the Study:
- To develop a novel, simple, and automated methodology for detecting height outliers in pediatric EHR data.
- To enhance the reliability of growth assessment tools and body proportionality indices.
Main Methods:
- Constructed two independent cohorts (training and testing) of children aged 2-8 years.
- Utilized monotonic Bayesian additive regression trees to predict height based on age, gender, race, and weight.
- Validated the model's performance in identifying height outliers using a manually reviewed testing cohort.
Main Results:
- The model explained 82.2% (training) and 75.3% (testing) of height variation (R²).
- Achieved excellent discriminatory ability for outlier detection in the testing cohort (Area Under the Curve = 0.841).
- Demonstrated good sensitivity (0.713) and specificity (0.793) with a cutoff of 0.075.
Conclusions:
- Developed a reliable and largely automated method for identifying height outliers in pediatric EHR.
- This methodology improves the accuracy of height data, ensuring reliable body proportionality indices like BMI.
- Facilitates better assessment of children's growth trajectories and health status.
Objective:
To create a new methodology that has a single simple rule to identify height outliers in the electronic health records (EHR) of children.
Methods:
We constructed 2 independent cohorts of children 2 to 8 years old to train and validate a model predicting heights from age, gender, race and weight with monotonic Bayesian additive regression trees. The training cohort consisted of 1376 children where outliers were unknown. The testing cohort consisted of 318 patients that were manually reviewed retrospectively to identify height outliers.
Results:
The amount of variation explained in height values by our model, R2 , was 82.2% and 75.3% in the training and testing cohorts, respectively. The discriminatory ability to assess height outliers in the testing cohort as assessed by the area under the receiver operating characteristic curve was excellent, 0.841. Based on a relatively aggressive cutoff of 0.075, the outlier sensitivity is 0.713, the specificity 0.793; the positive predictive value 0.615 and the negative predictive value is 0.856.
Conclusions:
We have developed a new reliable, largely automated, outlier detection method which is applicable to the identification of height outliers in the pediatric EHR. This methodology can be applied to assess the veracity of height measurements ensuring reliable indices of body proportionality such as body mass index.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Outliers and Influential Points
Regression Toward the Mean
Polygenic Traits
What Are Outliers?
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Variation: Normal Distribution, Range, and Standard Deviation

