Related Experiment Video
Updated: May 30, 2026

Using a Classroom-Based Deese Roediger McDermott Paradigm to Assess the Effects of Imagery on False Memories
Published on: November 14, 2018
Can the false-discovery rate be misleading?
Rodrigo Barboza1, Daniel Cociorva, Tao Xu
1Systems Engineering and Computer Science Program, COPPE, Federal University of Rio de Janeiro, Rio de Janeiro, Brazil.
The standard decoy-database approach in shotgun proteomics may overestimate confidence due to overfitting. A new semi-labeled decoy method statistically identifies and corrects for overfitted results, ensuring reliable protein identification.
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- The decoy-database approach is the standard for evaluating confidence in shotgun proteomic identification.
- This method uses decoy (random) sequences to estimate the false-discovery rate (FDR).
- Overfitting can occur, leading to inflated identification numbers that do not reflect the true FDR.
Purpose of the Study:
- To demonstrate that the decoy-database approach can be susceptible to overfitting.
- To introduce a novel method, the semi-labeled decoy approach, to detect and address overfitting.
- To improve the reliability of protein identifications in shotgun proteomics.
Main Methods:
- Implementation of the standard decoy-database approach in shotgun proteomic data analysis.
- Development and application of a modified semi-labeled decoy strategy.
- Statistical analysis to determine the presence and extent of overfitting.
Main Results:
- The study reveals that apparent good results from the decoy-database approach can be artifacts of overfitting.
- Overfitting inflates protein identification numbers, misrepresenting the actual false-discovery rate.
- The semi-labeled decoy approach successfully identifies overfitted results.
Conclusions:
- The conventional decoy-database approach may yield misleadingly high confidence scores due to overfitting.
- The semi-labeled decoy approach provides a statistically sound method to detect and mitigate overfitting.
- This advancement enhances the accuracy and reliability of protein identifications in shotgun proteomics.
Related Concept Videos
False Memories
One primary source of false memories is misattribution, where individuals incorrectly associate external information with...
Understanding Deception
Errors In Hypothesis Tests
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
