Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

5.9K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.9K
Modern Molecular Taxonomy01:29

Modern Molecular Taxonomy

58
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
58
Protein Families02:47

Protein Families

15.5K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
15.5K
Conserved Binding Sites01:49

Conserved Binding Sites

4.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.3K
Gene Evolution - Fast or Slow?02:05

Gene Evolution - Fast or Slow?

7.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
7.2K
Catalytically Perfect Enzymes01:07

Catalytically Perfect Enzymes

4.0K
The theory of catalytically perfect enzymes was first proposed by W.J. Albery and J. R. Knowles in 1976. These enzymes catalyze biochemical reactions at high-speed. Their catalytic efficiency values range from 108-109 M-1s-1. These enzymes are also called 'diffusion-controlled' as the only rate-limiting step in the catalysis is that of the substrate diffusion into the active site. Examples include triose phosphate isomerase, fumarase, and superoxide dismutase.
 
Most enzymes...
4.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Correction to "AstraMEV (AI-Guided Structural Assembly of Multi-Epitope Vaccines) Against Infectious Bronchitis Virus".

Journal of chemical information and modeling·2026
Same author

Rapid Affinity Characterization of Cell-Free Expressed Nanobodies Directly in Lysate with Biolayer Interferometry via Antigen-Immobilized Biosensors.

ACS synthetic biology·2026
Same author

Improving crash data quality by identifying misclassified alcohol-involved crashes using NLP on narrative data.

Journal of safety research·2026
Same author

Dynamic, Unconstrained Optimization of Secreted Enzyme Production in Fed-Batch Fermentation Using Reinforcement Learning.

Biotechnology and bioengineering·2026
Same author

High-Throughput FRET Affinity Screening Technique (HTFAST) For Cell-Free Expressed Binding Protein Characterization.

bioRxiv : the preprint server for biology·2026
Same author

SMART sensors for cell growth monitoring in closed system G-Rex for scalable, cost-conscious cell and gene therapy manufacturing.

Cytotherapy·2026

Related Experiment Video

Updated: Jul 29, 2025

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

68.8K

Effects of Sequence Features on Machine-Learned Enzyme Classification Fidelity.

Sakib Ferdous1, Ibne Farabi Shihab2, Nigel F Reuel1

  • 1Department of Chemical and Biological Engineering, Iowa State University.

Biochemical Engineering Journal
|May 22, 2023
PubMed
Summary

Enzyme Commission (EC) number assignment algorithms perform best with enzymes 300-500 amino acids long. Performance varies by enzyme class, with translocases being easiest to predict and hydrolases/oxidoreductases most challenging.

More Related Videos

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
09:34

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data

Published on: September 25, 2021

4.0K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Related Experiment Videos

Last Updated: Jul 29, 2025

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

68.8K
A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
09:34

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data

Published on: September 25, 2021

4.0K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Area of Science:

  • Bioinformatics
  • Enzymology
  • Computational Biology

Background:

  • Enzyme Commission (EC) number assignment is crucial for enzyme classification.
  • Current algorithms utilize statistics, homology, and machine learning based on sequence information.
  • Benchmarking these algorithms against sequence features is needed for optimizing enzyme design.

Purpose of the Study:

  • To benchmark the performance of enzyme classification algorithms based on sequence features.
  • To identify optimal classification windows for *de novo* enzyme sequence generation and design.
  • To develop efficient workflows for processing large enzyme sequence datasets.

Main Methods:

  • Developed parallelization and visualization workflows to process >500,000 annotated sequences.
  • Applied workflows to the SwissProt database using ECpred, DeepEC, Deepre, and BENZ-ws classifiers.
  • Analyzed algorithm performance as a function of enzyme chain length and amino acid composition (AAC).

Main Results:

  • All classifiers showed peak performance for enzymes between 300 and 500 amino acids in length.
  • Translocases (EC-6) were predicted most accurately; hydrolases (EC-3) and oxidoreductases (EC-1) were least accurate.
  • Classifiers performed best within common amino acid composition ranges and ECpred demonstrated high consistency.

Conclusions:

  • Enzyme length and amino acid composition significantly impact EC number assignment accuracy.
  • Optimal design spaces for synthetic enzyme generation can be identified using these benchmarking workflows.
  • The developed workflows facilitate the evaluation of new enzyme classification algorithms.