Related Experiment Video
Updated: Nov 28, 2025

Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
Essential gene prediction using limited gene essentiality information-An integrative semi-supervised machine learning
Sutanu Nandi1,2, Piyali Ganguli1,2, Ram Rup Sarkar1,2
1Chemical Engineering and Process Development, CSIR-National Chemical Laboratory, Pune, Maharashtra, India.
This study introduces a new machine learning pipeline for predicting essential genes, crucial for organism survival, especially in understudied pathogens. The method accurately identifies vital genes even with minimal experimental data, aiding in drug discovery.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Essential gene prediction identifies genes critical for organism survival.
- Machine learning (ML) aids gene essentiality prediction, but current methods struggle with limited experimental data.
- Less-explored, disease-causing organisms often lack sufficient data for accurate essential gene annotation.
Purpose of the Study:
- To develop a novel ML pipeline for predicting essential genes in organisms with limited experimental data.
- To improve the annotation of essential genes for understudied, pathogenic organisms.
- To provide a robust strategy for identifying potential therapeutic targets.
Main Methods:
- A pipeline combining unsupervised feature selection, Kamada-Kawai dimension reduction, and Laplacian Support Vector Machine (LapSVM) was developed.
- A novel Semi-Supervised Model Selection Score, analogous to auROC, was proposed for model selection with limited data.
- Genome-scale metabolic networks were utilized for predicting essential and non-essential genes.
Main Results:
- The pipeline achieved high accuracy (auROC > 0.85) in essential gene prediction, even with as little as 1% labeled data.
- Unsupervised feature selection and dimension reduction revealed distinct clustering patterns for essential and non-essential genes.
- The method demonstrated high accuracy and universality across Eukaryotes and Prokaryotes, including Leishmania sp.
Conclusions:
- The proposed ML pipeline effectively predicts essential genes using limited labeled data, addressing a critical gap in bioinformatics.
- This approach facilitates essential gene annotation for poorly characterized organisms, aiding in the identification of novel therapeutic targets.
- The strategy offers a valuable tool for antibiotic and vaccine development against parasitic diseases.
Related Concept Videos
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
Gene Evolution - Fast or Slow?
In contrast, regions which code...

