Related Experiment Video
Updated: Jun 21, 2026

11:59
Competitive Genomic Screens of Barcoded Yeast Libraries
Published on: August 11, 2011
Methods for labeling error detection in microarrays based on the effect of data perturbation on the regression model
Chen Zhang1, Chunguo Wu, Enrico Blanzieri
1College of Computer Science and Technology, Jilin University, 130012 China.
Bioinformatics (Oxford, England)
|August 8, 2009
Summary
This study introduces new methods to detect mislabeled microarray samples by measuring data perturbation effects. The proposed algorithms, particularly PRAPIV, improve accuracy and recall over existing techniques.
Area of Science:
- Bioinformatics
- Machine Learning
- Genomics
Background:
- Mislabeled samples in gene expression profiles hinder supervised learning due to disease subtype similarity and misdiagnosis.
- Existing methods like LOOE-sensitivity have limitations in accurately measuring data perturbation effects for mislabeled sample detection.
Purpose of the Study:
- To design a novel method for detecting mislabeled microarray samples by effectively utilizing data perturbation measurements.
- To address the performance limitations of previous data perturbation-based algorithms.
Main Methods:
- Defined a perturbing influence value (PIV) using a support vector machine (SVM) regression model to quantify data perturbation effects.
- Developed Column Algorithm (CAPIV), Row Algorithm (RAPIV), and progressive Row Algorithm (PRAPIV) based on the PIV for mislabeled sample detection.
Main Results:
- All proposed methods (CAPIV, RAPIV, PRAPIV) outperformed the LOOE-sensitivity algorithm on artificial and microarray datasets.
- The progressive Row Algorithm (PRAPIV) demonstrated superior precision and high recall compared to simple SVM and CL-stability.
Conclusions:
- The developed PIV-based algorithms offer a more effective approach to detecting mislabeled samples in microarray data.
- PRAPIV shows significant improvements in detection accuracy, making it a valuable tool for data cleaning in bioinformatics.
Related Concept Videos
Types of Errors: Detection and Minimization
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Systematic Error: Methodological and Sampling Errors
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least squares (OLS)...

