Related Experiment Video
Updated: Oct 30, 2025

Sequencing of Plant Wall Heteroxylans Using Enzymic, Chemical Methylation and Physical Mass Spectrometry, Nuclear Magnetic Resonance Techniques
Published on: March 24, 2016
Machine learning reveals sequence-function relationships in family 7 glycoside hydrolases
Japheth E Gado1, Brent E Harrison2, Mats Sandgren3
1Department of Chemical and Materials Engineering, University of Kentucky, Lexington, Kentucky, USA; Renewable Resources and Enabling Sciences Center, National Renewable Energy Laboratory, Golden, Colorado, USA.
Machine learning accurately predicts glycoside hydrolase 7 (GH7) enzyme function based on active-site loop lengths and specific sequence positions. This data-driven approach aids in understanding and engineering these crucial cellulose-degrading enzymes.
Area of Science:
- Biochemistry
- Enzymology
- Computational Biology
Background:
- Family 7 glycoside hydrolases (GH7) are key enzymes in cellulose degradation, crucial for natural and industrial processes.
- GH7 enzymes are often bimodular, featuring catalytic and carbohydrate-binding domains (CBM), with active sites binding cello-oligomers.
- Understanding the sequence-structure-function relationships in GH7 enzymes, including cellobiohydrolases (CBH) and endoglucanases (EG), remains incomplete.
Purpose of the Study:
- To apply machine learning for data-driven insights into GH7 enzyme sequence, structure, and function.
- To identify key sequence features correlating with GH7 functional subtypes (CBH vs. EG) and CBM presence.
- To explore novel sequence positions potentially influencing GH7 functional variation.
Main Methods:
- Machine learning models were trained using active-site loop residue counts to classify GH7 subtypes.
- Classification rules were derived based on specific residue positions predicting functional subtype.
- A random forest model analyzed catalytic domain residues to predict CBM presence.
Main Results:
- Machine learning models accurately discriminated between GH7 CBHs and EGs (up to 99% accuracy) based on active-site loop lengths (A4, B2, B3, B4).
- Specific residues at 42 sequence positions predicted functional subtype with over 87% accuracy.
- A random forest model predicted CBM presence with 89.5% accuracy using 19 catalytic domain positions.
Conclusions:
- Active-site loop lengths and specific sequence positions are strong predictors of GH7 functional subtypes and CBM presence.
- Machine learning successfully recapitulates and expands upon experimentally identified key sequence positions for GH7 activity.
- These findings offer a foundation for understanding and engineering GH7 enzymes for improved cellulose degradation.
Related Concept Videos
Protein Families
Protein Families
Oligosaccharide Assembly
Multiple sugar molecules that may or may...
Lysosomal Hydrolases
Hydrolysis
Hydrolysis is a chemical reaction in which the addition of water breaks down a polymer into its simpler monomer units. For example, peptides break into amino acids, carbohydrates into simple sugars, and DNA into nucleotides. Enzymes often facilitate these processes.
Hydrolysis Reverses Dehydration Synthesis
Complex carbohydrates can be broken down by breaking the bonds between individual sugar units. The reaction breaks a glycosidic bond as water is added to the compound. The...
Biosynthesis of Polysaccharides

