Related Experiment Video
Updated: Aug 3, 2025

08:10
Multi-enzyme Screening Using a High-throughput Genetic Enzyme Screening System
Published on: August 8, 2016
8.9K
Prediction of enzymatic function with high efficiency and a reduced number of features using genetic algorithm
Diogo R Reis1, Bruno C Santos1, Lucas Bleicher2
1Pontifical Catholic University of Minas Gerais - PUC Minas, 500 Dom José Gaspar Street, Building 20, Coração Eucarístico, Belo Horizonte, MG 30535-901, Brazil.
Computers in Biology and Medicine
|April 7, 2023
Summary
This study uses machine learning and genetic algorithms to identify key protein characteristics for enzyme classification. A novel approach significantly reduced features while improving prediction accuracy, making protein function identification more efficient.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- The post-genomic era necessitates efficient methods for protein function identification.
- Machine learning approaches, particularly feature-based methods, are crucial in bioinformatics.
- Understanding protein structures (primary, secondary, tertiary, quaternary) is key to function prediction.
Purpose of the Study:
- To investigate protein characteristics that enhance machine learning model quality for enzyme class prediction.
- To apply dimensionality reduction techniques and Support Vector Machine (SVM) classification.
- To develop and evaluate feature selection methods, including a novel genetic algorithm approach.
Main Methods:
- Utilized Factor Analysis for feature extraction/transformation.
- Developed a multi-objective genetic algorithm for feature selection.
- Compared the proposed genetic algorithm with other feature selection methods.
- Employed Support Vector Machine (SVM) for enzyme classification.
Main Results:
- A feature subset generated by the multi-objective genetic algorithm reduced the dataset by approximately 87%.
- This subset achieved an F-measure of 85.78%, significantly improving classification model quality.
- A reduced set of 28 features achieved over 80% F-measure for four out of six enzyme classes.
Conclusions:
- Satisfactory enzyme classification performance can be achieved using a significantly reduced set of protein characteristics.
- The developed multi-objective genetic algorithm is effective for selecting relevant features in bioinformatics.
- Openly available datasets and implementations facilitate further research in protein function prediction.

