Related Experiment Videos
Machine Learning in Therapeutic Research: The Hard Work of Outlier Detection in Large Data
1European College Pharmaceutical Medicine, Lyon, France c/o Department of Medicine, Albert Schweitzer Hospital, Dordrecht, Netherlands.
American Journal of Therapeutics
|February 23, 2013
Summary
Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) effectively identifies outliers in large datasets, outperforming traditional regression analysis in therapeutic research. This advanced clustering method reveals previously unrecognized patient groups, crucial for accurate data interpretation.
Area of Science:
- Data Science
- Biostatistics
- Machine Learning
Background:
- Outlier detection in large datasets is challenging with traditional methods like data plots and regression lines.
- The number of outliers often increases linearly with dataset size, necessitating advanced analytical approaches.
- Unidentified outliers can significantly impact research findings and lead to erroneous conclusions in therapeutic studies.
Purpose of the Study:
- To evaluate the efficacy of Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) for detecting previously unrecognized outliers in large datasets.
- To compare the performance of BIRCH clustering against traditional regression analysis in outlier identification.
- To demonstrate the utility of BIRCH for analyzing complex patient data in therapeutic research.
Main Methods:
- Utilized a simulated dataset and a real-world dataset (iatrogenic admissions) for analysis.
- Employed SPSS statistical software for data analysis.
- Applied BIRCH clustering algorithm to identify data points that do not conform to established clusters.
Main Results:
- In a dataset of 50 mentally depressed persons, regression analysis failed to detect outliers, while BIRCH identified a relevant outlier cluster of 7 patients (14%).
- In a dataset of 576 iatrogenic admissions, BIRCH analysis revealed an outlier cluster of 174 patients (30%) with an extremely high number of comedications, a finding not predicted by loglinear analysis.
- BIRCH successfully identified significant outlier groups in both simulated and real-world datasets.
Conclusions:
- Systematic outlier assessment is critical in large-scale therapeutic research to prevent potentially catastrophic consequences.
- Traditional methods like regression analysis are insufficient for detecting outliers in complex, large datasets.
- BIRCH clustering offers a robust and suitable method for identifying relevant outlier clusters, particularly in large datasets, enhancing data interpretation and research validity.
Related Concept Videos
Outliers and Influential Points
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the vertical...
What Are Outliers?
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Quantifying and Rejecting Outliers: The Grubbs Test
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Regression Toward the Mean
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Steps in Outbreak Investigation
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis: