cpiVAE: Robust and Interpretable Cross-Platform Proteomics Imputation.
Yuxiang Li1,2, ThuyVy Duong3, Mary R Rooney4
1Department of Biomedical Engineering, Johns Hopkins University, Baltimore, Maryland, USA.
Discordant plasma proteomic data across platforms hinders research. A new AI model, cross-platform proteomics imputation variational autoencoder (cpiVAE), accurately imputes protein levels, improving data integration for powerful meta-analyses.
Area of Science:
- Biochemistry
- Bioinformatics
- Genomics
Background:
- Plasma proteomic studies frequently employ diverse high-throughput platforms, leading to discordant measurements for identical proteins.
- This data discordance impedes effective cross-study integration and limits the power of meta-analyses and biomarker discovery.
- Integrating proteomics data is crucial for enhancing statistical power and deepening the understanding of proteome-phenotype relationships.
Purpose of the Study:
- To develop a deep generative model for bidirectional imputation of protein abundances between Olink and SomaScan platforms.
- To establish a method that improves cross-platform proteomics data integration and enables more powerful downstream analyses.
- To provide an interpretable and generalizable solution for harmonizing proteomics data from different experimental platforms.
Main Methods:
- Development of a cross-platform proteomics imputation variational autoencoder (cpiVAE), a deep generative model.
- Training the cpiVAE model using paired plasma proteomic measurements from the China Kadoorie Biobank (CKB) cohort.
- Evaluating cpiVAE performance against established methods like k-nearest neighbors (KNN) and Weighted Nearest Neighbors (WNN) on independent datasets.
Main Results:
- cpiVAE achieved up to 30% higher correlation between imputed and true protein values compared to KNN and WNN benchmarks.
- The model demonstrated strong generalization to an independent Atherosclerosis Risk in Communities Study (ARIC) cohort without retraining.
- Imputed protein levels accurately reflected associations with clinical phenotypes and enhanced statistical power in meta-analysis simulations.
- A post-hoc feature importance matrix provided interpretability, with extracted protein pair features overlapping with known biological interactions in the STRING database.
Conclusions:
- cpiVAE offers an accurate, generalizable, and interpretable solution for cross-platform proteomics imputation.
- The framework facilitates integrated analyses across studies utilizing different proteomics measurement platforms.
- Open-source availability of the cpiVAE framework and pre-trained model weights promotes wider adoption and data integration in proteomic research.
More Related Videos
07:01Navigating the Mass Spectrometry-Based Proteomic Data Using Free Computational Tools
Published on: August 19, 2025
09:52A Clinical Metaproteomics Workflow Implemented within Galaxy Bioinformatics Platform to Analyze Host-Microbiome Interactions Underlying Human Disease
Published on: January 10, 2025
Related Concept Videos
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Improving Translational Accuracy
Improving Translational Accuracy
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
