Related Experiment Video
Updated: May 3, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.0K
Training set selection for the prediction of essential genes
Jian Cheng1, Zhao Xu2, Wenwu Wu3
1College of Life Sciences and State Key Laboratory of Crop Stress Biology in Arid Areas, Northwest A&F University, Yangling, Shaanxi, China ; Bioinformatics Center, Northwest A&F University, Yangling, Shaanxi, China.
Plos One
|January 28, 2014
Summary
Selecting reliable training sets is crucial for accurately predicting essential genes in microorganisms. Integrated training sets, based on specific criteria, significantly improve prediction accuracy and stability for genome-wide essential gene identification.
Area of Science:
- Genomics
- Computational Biology
- Systems Biology
Background:
- Computational models exist for transferring gene essentiality annotations between organisms.
- Predicting essential genes in poorly-studied microorganisms is challenging due to difficulties in selecting appropriate training sets.
Purpose of the Study:
- To apply a machine learning approach to predict essential genes in 21 microorganisms.
- To determine optimal criteria for selecting training sets to improve prediction accuracy.
- To evaluate the performance of incomplete versus integrated training sets.
Main Methods:
- Reciprocal prediction of essential genes using a machine learning approach across 21 microorganisms.
- Development and application of four criteria for training set selection: reliability, consistent growth conditions, phylogenetic relatedness, and similar phenotypes/lifestyles.
- Analysis of training set size (minimum 10% of total genes) and comparison of single vs. integrated training sets.
Main Results:
- Training set selection significantly impacts the accuracy of essential gene prediction.
- Integrated training sets demonstrated superior stability and accuracy compared to single-organism sets.
- Rational selection of training sets based on established criteria outperformed random selection.
Conclusions:
- Empirical guidance is provided for selecting training sets to enhance genome-wide essential gene identification.
- The study highlights the importance of training set quality and composition for accurate cross-species gene essentiality predictions.

