Related Experiment Videos
Relabeling algorithm for retrieval of noisy instances and improving prediction quality.
1Health Systems Management, Rush University Medical Center, 1700 W. Van Buren Street, Chicago, IL 60612, USA.
Computers in Biology and Medicine
|January 26, 2010
Summary
This study introduces a novel relabeling algorithm to improve prediction quality by identifying and correcting noisy data instances. The algorithm enhances knowledge generalization and confidence, proving effective across diverse datasets.
Area of Science:
- Machine Learning
- Data Mining
- Predictive Analytics
Background:
- Noisy data instances with binary outcomes can significantly degrade prediction quality in machine learning models.
- Existing methods often prioritize classification accuracy, potentially overlooking crucial aspects like knowledge generalization and data confidence.
Purpose of the Study:
- To present a novel relabeling algorithm designed for the retrieval and correction of noisy instances in datasets with binary outcomes.
- To enhance prediction quality by focusing on knowledge generalization and confidence, rather than solely on classification accuracy.
Main Methods:
- The algorithm iteratively retrieves, selects, and relabels data instances, effectively transforming the decision space.
- A confidence index was developed, integrating classification accuracy, prediction error, dataset impurities, and cluster purities.
- The approach was validated on benchmark UCI datasets and bladder cancer immunotherapy data.
Main Results:
- A subset of stable instances (7%-51%) with high confidence (64%-99.44%) was identified, alongside the most noisy instances.
- The relabeled instances and their confidence indexes were validated by domain experts.
- The algorithm demonstrated successful application on diverse datasets, including medical and benchmark data.
Conclusions:
- The proposed relabeling algorithm effectively improves prediction quality by addressing noisy instances.
- The emphasis on knowledge generalization and confidence provides a robust measure of data reliability.
- The algorithm shows potential for broad applicability across medical, industrial, and service domains with modifications.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Prediction Intervals
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...