Related Experiment Video
Updated: Nov 4, 2025

10:37
Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
Published on: November 15, 2017
12.3K
A community effort to identify and correct mislabeled samples in proteogenomic studies
Seungyeul Yoo1,2,3, Zhiao Shi4,5, Bo Wen4,5
1Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY 10029, USA.
Patterns (New York, N.Y.)
|May 26, 2021
Summary
Sample mislabeling is a major issue in multi-omic research. A challenge led to COSMO, an open-source tool that accurately identifies and corrects these errors in proteogenomic data.
Area of Science:
- Genomics
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- Sample mislabeling and misannotation are persistent challenges in scientific research.
- These errors are particularly common in complex, large-scale multi-omic studies.
- Automated quality control methods are urgently needed to detect and correct sample mix-ups in multi-omic data.
Purpose of the Study:
- To establish a framework for evaluating methods to identify and correct sample mislabels in integrative proteogenomic studies.
- To benchmark the performance of various computational approaches for detecting sample mislabeling.
- To foster the development of robust solutions for multi-omic data quality control.
Main Methods:
- A crowdsourced challenge (precisionFDA NCI-CPTAC Multi-omics Enabled Sample Mislabeling Correction Challenge) was organized.
- Data scientists submitted diverse methods for mislabel identification and correction.
- Performance of submitted methods was systematically evaluated on simulated and real multi-omic datasets.
Main Results:
- The challenge attracted numerous submissions from global participants.
- Performance varied significantly across the submitted mislabeling correction methods.
- Post-challenge collaboration resulted in the development of the open-source software COSMO.
Conclusions:
- COSMO demonstrates high accuracy and robustness in identifying and correcting sample mislabels.
- The developed software provides a valuable tool for quality control in proteogenomic studies.
- This work highlights the effectiveness of crowdsourced challenges in advancing multi-omic data integrity.

