Related Experiment Video
Updated: Mar 25, 2026

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
9.7K
Enabling cross-indication protein expression analysis using a curated pan-cancer dataset and a tailored workflow
Jixin Wang1, Xiaowen Tian2, Wen Yu3
1Oncology Data Science & AI, AstraZeneca, Gaithersburg, MD, USA.
Scientific Reports
|March 24, 2026
Summary
We developed a robust method to normalize proteomic data from the National Cancer Institute's Clinical Proteomic Tumor Analysis Consortium (CPTAC) pan-cancer study. This curated dataset enables reliable cross-cohort protein expression analysis for cancer research.
Area of Science:
- Proteomics
- Cancer Biology
- Bioinformatics
Background:
- The National Cancer Institute's Clinical Proteomic Tumor Analysis Consortium (CPTAC) generated comprehensive multi-omics data for over 1,000 tumors.
- Comparing protein expression across CPTAC cohorts is difficult due to missing data and varied expression patterns.
Purpose of the Study:
- To create a curated and normalized pan-cancer protein expression dataset from CPTAC data.
- To enable robust cross-cohort protein expression analysis for the cancer research community.
Main Methods:
- Developed a novel algorithm for selecting robustly expressed proteins within CPTAC cohorts.
- Applied a cohort hybrid imputation approach for protein abundance values.
- Utilized intensity-based absolute quantification and compared global vs. smooth quantile normalization.
Main Results:
- Global quantile normalization showed superior performance compared to smooth quantile normalization and no normalization.
- Higher rank correlation across cancer cohorts was observed between CPTAC and TCGA using global quantile normalization.
- The proposed workflow effectively addresses missing data and normalizes protein expression patterns.
Conclusions:
- Combining cohort hybrid imputation with global quantile normalization creates an effective normalized CPTAC pan-cancer protein dataset.
- This normalized dataset will facilitate the study of protein expression across diverse cancer types.
- The findings accelerate pan-cancer discovery research by enabling reliable proteomic data comparison.

