Related Experiment Video
Updated: Sep 6, 2025

04:40
Tactile Semiautomatic Passive-Finger Angle Stimulator TSPAS
Published on: July 30, 2020
3.0K
Nonparametric percentile curve estimation for a nonnegative marker with excessive zeros
Oke Gerke1,2, Robyn L McClelland3
1Department of Nuclear Medicine, Odense University Hospital, Odense, Denmark.
Methodsx
|July 5, 2022
Summary
This study introduces a new nonparametric method for estimating percentile curves for medical markers like the Agatston score. The approach accurately models skewed data, including zero-inflated scores, crucial for preventive medicine and risk assessment.
Area of Science:
- Biostatistics
- Preventive Medicine
- Medical Imaging Analysis
Background:
- Percentile curves are established for infant growth metrics.
- The Agatston score, introduced in 1990, is a key marker in preventive cardiology.
- Non-negative scores like the Agatston score often exhibit an excess of zero values in healthy populations.
Purpose of the Study:
- To develop a nonparametric method for estimating percentile curves for medical markers, particularly those with zero-inflated distributions.
- To address the challenge of accurately representing skewed data, such as the Agatston score, in percentile curve analysis.
Main Methods:
- A nonparametric approach using lowess smoothing for marker-positive scores.
- Transposing smoothed curves based on estimated proportions of zero scores.
- Validation through a simulation study with varying sample sizes (N=1,000 to 10,000).
Main Results:
- The method accurately estimates 50th, 75th, and 90th percentile curves, even with exponentially distributed markers and increasing zero proportions.
- Demonstrated robustness against outliers and adherence to the non-crossing property of percentile curves.
- Successfully applied to subgroup data, showcasing applicability to highly skewed datasets.
Conclusions:
- The developed nonparametric method provides an explicit, transferable, and reproducible procedure for percentile curve estimation.
- This approach is versatile and applicable to a broad range of markers and scores across diverse scientific disciplines beyond cardiovascular medicine.
- The method effectively handles zero-inflated and skewed data, enhancing its utility in medical research and clinical practice.
More Related Videos
Related Concept Videos
Percentile
7.0K
A percentile indicates the relative standing of a data value when data are sorted into numerical order from smallest to largest. It represents the percentages of data values that are less than or equal to the pth percentile. For example, 15% of data values are less than or equal to the 15th percentile.
7.0K
What Are Outliers?
4.2K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
4.2K
Outliers and Influential Points
4.2K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.2K
z Scores and Area Under the Curve
11.3K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
11.3K
Modified Boxplots
10.0K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
10.0K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K

