Related Experiment Video
Updated: Feb 24, 2026

09:39
Mapping Dysfunctional Protein-Protein Interactions in Disease
Published on: October 24, 2025
925
Unsupervised learning reveals novel disease-associated proteins in high-dimensional human proteomic data
Elvis Bernard1, Yiling Wang2, Manlin Chen2
1School of Environmental Science and Engineering, Hainan University, Haikou, 570228, China. elvis.bernard@hainanu.edu.cn.
Scientific Reports
|February 22, 2026
Summary
A new framework, DIRAM/COD, analyzes large proteomic datasets by combining dimensionality reduction and unsupervised learning. This approach identifies known and novel disease biomarkers, advancing precision medicine and biomarker discovery.
Area of Science:
- Biomedical Data Science
- Proteomics
- Computational Biology
Background:
- Precision medicine generates massive proteomic datasets, posing analytical challenges.
- Supervised learning is common but may miss subtle patterns.
- Unsupervised learning can reveal hidden relationships but struggles with high dimensionality.
Purpose of the Study:
- To develop a novel computational framework for analyzing high-dimensional proteomic data.
- To address the limitations of traditional supervised and unsupervised learning methods in large-scale proteomic analysis.
Main Methods:
- Developed the Dimensionality Reduction with Avoidance of Missing/COmmunity Detection (DIRAM/COD) framework.
- Combined dimensionality reduction techniques with unsupervised learning.
- Applied the framework to the UK Biobank proteomic dataset (2,923 proteins, 52,691 participants).
Main Results:
- Successfully analyzed a large-scale proteomic dataset.
- Confirmed established biomarkers for hypertension (UBE2L6) and leukemia (LRCH4).
- Identified novel protein candidates, including IGF2BP3 for celiac disease.
Conclusions:
- The DIRAM/COD framework effectively analyzes high-dimensional proteomic data.
- This approach facilitates the discovery of both known and novel disease biomarkers.
- Opens new avenues for biomarker and therapeutic target identification in precision medicine.
Related Concept Videos
Proteomics
10.0K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
10.0K
Protein Networks
4.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.6K
Genome-wide Association Studies-GWAS
15.9K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.9K

