Related Experiment Video
Updated: Feb 9, 2026

Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
C-PUGP: A cluster-based positive unlabeled learning method for disease gene prediction and prioritization
Akram Vasighizaker1, Saeed Jalili1
1Computer Engineering Department, Tarbiat Modares University, Tehran, Iran.
This study introduces a novel machine learning approach for identifying disease genes using clustering and One-Class classification. The method improves candidate gene prioritization by creating a reliable negative dataset, achieving high precision and recall.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning in genomics
Background:
- Accurate disease gene detection is crucial for understanding disease mechanisms and developing treatments.
- Machine learning methods are widely used for identifying candidate disease genes, but often face challenges with insufficient negative data.
- Semi-supervised learning on positive and unlabeled data has shown promise but can be improved.
Purpose of the Study:
- To propose a novel Positive Unlabeled (PU) learning technique for more reliable disease gene identification.
- To address the challenge of limited negative data in candidate gene detection.
- To enhance the accuracy and performance of disease gene prioritization.
Main Methods:
- A new PU learning technique combining clustering and One-Class classification is proposed.
- A three-step process is used to generate a Reliable Negative (RN) set: clustering positive data, learning One-Class classifiers, and selecting the intersection of negative data.
- A Support Vector Machine (SVM) binary classifier is employed for candidate disease gene identification and ranking.
Main Results:
- The proposed method achieved high performance metrics: 92.8% precision, 93.6% recall, and 93.1% F-measure.
- Performance, particularly in F-measure, showed an 11.7% improvement compared to existing methods.
- A notable 6% increase in prioritization results was observed.
Conclusions:
- The novel PU learning technique effectively addresses the challenge of limited negative data in disease gene detection.
- The proposed method significantly outperforms existing approaches in identifying and ranking candidate disease genes.
- This approach offers a more accurate and reliable tool for genomic research in disease understanding.
Related Concept Videos
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Velocity and Position by Integral Method
Consider an example to calculate the velocity and position from the acceleration function. A motorboat is traveling at a constant velocity of 5.0 m/s when it starts to decelerate to arrive at the dock. Its acceleration is...
Velocity and Position by Graphical Method
Position-effect Variegation
Position of Equilibrium in Acid-Base Reactions

