Related Experiment Video
Updated: Jul 29, 2025

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Effects of Sequence Features on Machine-Learned Enzyme Classification Fidelity
Sakib Ferdous1, Ibne Farabi Shihab2, Nigel F Reuel1
1Department of Chemical and Biological Engineering, Iowa State University.
Abstract:
Assigning enzyme commission (EC) numbers using sequence information alone has been the subject of recent classification algorithms where statistics, homology and machine-learning based methods are used. This work benchmarks performance of a few of these algorithms as a function of sequence features such as chain length and amino acid composition (AAC). This enables determination of optimal classification windows for de novo sequence generation and enzyme design. In this work we developed a parallelization workflow which efficiently processes >500,000 annotated sequences through each candidate algorithm and a visualization workflow to observe the performance of the classifier over changing enzyme length, main EC class and AAC. We applied these workflows to the entire SwissProt database to date (n = 565245) using two, locally installable classifiers, ECpred and DeepEC, and collecting results from two other webserver-based tools, Deepre and BENZ-ws. It is observed that all the classifiers exhibit peak performance in the range of 300 to 500 amino acids in length. In terms of main EC class, classifiers were most accurate at predicting translocases (EC-6) and were least accurate in determining hydrolases (EC-3) and oxidoreductases (EC-1). We also identified AAC ranges that are most common in the annotated enzymes and found that all classifiers work best in this common range. Among the four classifiers, ECpred showed the best consistency in changing feature space. These workflows can be used to benchmark new algorithms as they are developed and find optimum design spaces for the generation of new, synthetic enzymes.
More Related Videos
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Protein Families
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Catalytically Perfect Enzymes
Most enzymes...