Related Experiment Video
Updated: Oct 4, 2025

09:10
A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
9.3K
SMAP is a pipeline for sample matching in proteogenomics.
Ling Li1, Mingming Niu2, Alyssa Erickson1
1Department of Biology, University of North Dakota, Grand Forks, ND, 58202, USA.
Nature Communications
|February 9, 2022
Summary
Sample mix-ups in proteogenomics studies are common. A new pipeline, Sample Matching in Proteogenomics (SMAP), uses mass spectrometry data to verify sample identity, ensuring accurate human disease research.
Area of Science:
- Genomics and Proteomics
- Human Disease Research
- Bioinformatics
Background:
- Proteogenomics integrates genomics and proteomics for deeper human disease understanding.
- Sample mix-ups are a significant challenge in complex proteogenomics workflows.
- Ensuring data integrity is crucial for reliable research findings.
Purpose of the Study:
- To develop and validate a computational pipeline for verifying sample identity in proteogenomics.
- To address the pervasive issue of sample mix-up in large-scale studies.
- To enhance the accuracy and reliability of proteogenomic data analysis.
Main Methods:
- Developed Sample Matching in Proteogenomics (SMAP) pipeline.
- Inferred sample-specific protein-coding variants from quantitative mass spectrometry (MS) data.
- Aligned proteomic and genomic samples using two discriminant scores.
Main Results:
- SMAP accurately matches proteomic and genomic samples with ≥20% genotype data.
- Applied to the PsychENCODE BrainGVEX dataset, SMAP corrected 54 samples (19%).
- Corrections were validated using ribosome profiling and ATAC-seq data.
Conclusions:
- SMAP is an effective tool for sample verification in large-scale MS-based proteogenomics.
- The pipeline ensures data integrity, crucial for advancing human disease research.
- SMAP is publicly available, facilitating its adoption in the scientific community.

