Novel Pediatric Height Outlier Detection Methodology for Electronic Health Records via Machine Learning With

Rodney A Sparapani1, Bi Q Teng1, Julia Hilbrands2

  • 1From the Division of Biostatistics, Medical College of Wisconsin, Milwaukee, WI.

Insights

A new automated method reliably identifies height outliers in children's electronic health records (EHR). This approach ensures accurate body proportionality measurements, improving pediatric health assessments.

Area of Science:

  • Pediatric Endocrinology
  • Biostatistics
  • Health Informatics

Background:

  • Accurate height measurements are crucial for monitoring child growth and development.
  • Identifying outliers in electronic health records (EHR) is essential for data integrity.

Purpose of the Study:

  • To develop a novel, simple, and automated methodology for detecting height outliers in pediatric EHR data.
  • To enhance the reliability of growth assessment tools and body proportionality indices.

Main Methods:

  • Constructed two independent cohorts (training and testing) of children aged 2-8 years.
  • Utilized monotonic Bayesian additive regression trees to predict height based on age, gender, race, and weight.
  • Validated the model's performance in identifying height outliers using a manually reviewed testing cohort.

Main Results:

  • The model explained 82.2% (training) and 75.3% (testing) of height variation (R²).
  • Achieved excellent discriminatory ability for outlier detection in the testing cohort (Area Under the Curve = 0.841).
  • Demonstrated good sensitivity (0.713) and specificity (0.793) with a cutoff of 0.075.

Conclusions:

  • Developed a reliable and largely automated method for identifying height outliers in pediatric EHR.
  • This methodology improves the accuracy of height data, ensuring reliable body proportionality indices like BMI.
  • Facilitates better assessment of children's growth trajectories and health status.
Abstract

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.2K
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.3K
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Polygenic Traits01:18

Polygenic Traits

When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
66.7K
What Are Outliers?01:12

What Are Outliers?

Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
4.3K
Variation: Normal Distribution, Range, and Standard Deviation02:32

Variation: Normal Distribution, Range, and Standard Deviation

In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data, such as ethnicity, can be tabulated into a frequency count to provide information about the proportion, as well as the variety of groups in a sample or population. On the other hand, researchers can perform a wider set of calculations on quantitative data. The mean, mode, and median, for instance, are central tendency measures to identify a...
22.6K