Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Intelligent tool orchestration for rapid mechanistic model prototyping: MCP servers as AI-biology interfaces.

NPJ systems biology and applications·2026
Same author

IMPaCT-Data: A Federated Precision Medicine Infrastructure Associated with Science and Technology in Spain.

Studies in health technology and informatics·2026
Same author

Commonalities in frailty and psychopathology predict chronotype across severe mental disorders from a comorbidity perspective.

Psychological medicine·2026
Same author

Therapeutic subtypes of knee osteoarthritis: differential treatment effects among predicted endotypes in past clinical trials.

Arthritis research & therapy·2026
Same author

Leveraging training expertise to build capacity in computational personalised medicine.

Bioinformatics advances·2026
Same author

Deep representation learning for temporal inference in cancer omics: a systematic literature review.

Briefings in bioinformatics·2026

Related Experiment Video

Updated: Jun 26, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
06:50

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions

Published on: January 26, 2024

Automated alphabet reduction for protein datasets.

Jaume Bacardit1, Michael Stout, Jonathan D Hirst

  • 1ASAP research group, School of Computer Science, University of Nottingham, Jubilee Campus, Wollaton Road, Nottingham, NG8 1BB, UK. jaume.bacardit@nottingham.ac.uk

BMC Bioinformatics
|January 8, 2009
PubMed
Summary

We developed an automated method to reduce protein alphabet size for faster analysis without losing accuracy. This technique creates compact, informative protein representations for improved machine learning and data mining in structural bioinformatics.

More Related Videos

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
05:08

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins

Published on: July 8, 2025

An Optimized Quantitative Pull-Down Analysis of RNA-Binding Proteins Using Short Biotinylated RNA
07:55

An Optimized Quantitative Pull-Down Analysis of RNA-Binding Proteins Using Short Biotinylated RNA

Published on: February 17, 2023

Related Experiment Videos

Last Updated: Jun 26, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
06:50

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions

Published on: January 26, 2024

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
05:08

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins

Published on: July 8, 2025

An Optimized Quantitative Pull-Down Analysis of RNA-Binding Proteins Using Short Biotinylated RNA
07:55

An Optimized Quantitative Pull-Down Analysis of RNA-Binding Proteins Using Short Biotinylated RNA

Published on: February 17, 2023

Area of Science:

  • Structural Bioinformatics
  • Computational Biology
  • Machine Learning

Background:

  • Protein structure prediction relies on complex datasets.
  • Reducing alphabet size can accelerate machine learning and data mining.
  • Informative reduced alphabets aid in creating compact, human-friendly classification rules.

Purpose of the Study:

  • To develop an automated, robust protocol for protein alphabet reduction.
  • To create reduced alphabets that retain key biochemical information.
  • To enhance efficiency in structural bioinformatics applications.

Main Methods:

  • Utilized mutual information and advanced optimization techniques.
  • Applied the protocol to predict protein contact number and relative solvent accessibility.
  • Generated reduced alphabets of varying sizes (2, 3, 4, and 5 letters).

Main Results:

  • Five-letter alphabets achieved prediction accuracies comparable to the full amino acid alphabet.
  • The automated alphabets outperformed existing reduced alphabets from literature and human designs.
  • Extrapolation to evolutionary information-based protein representations showed minimal performance loss.

Conclusions:

  • The automated protocol effectively generates reduced alphabets for diverse protein datasets.
  • The process is domain-knowledge-free, relying on information theory.
  • Discovered novel amino acid groupings, offering new data interpretation perspectives.