Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key locations, protein...
Protein-protein Interfaces02:04

Protein-protein Interfaces

Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein-Protein Interfaces02:04

Protein-Protein Interfaces

Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Organization01:24

Protein Organization

Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Cracks in the AI Crystal Ball: Why Clinical Prediction Tools Fall Short in the Real World.

Journal of general internal medicine·2026
Same author

Exploring Gamification of the Incentive Spirometry Tracker Device: A Survey of Patients, Providers, and Healthcare Employees.

Anesthesiology·2026
Same author

Epigenome-Wide Association Study in Asian Cohort Identifies Novel DNA Methylation Markers for Carotid Intima-Media Thickness.

Research square·2026
Same author

Phase I trial of locoregional administration of autologous tumor-infiltrating lymphocytes in patients with uveal melanoma and liver metastases (the HAITILS trial).

Journal for immunotherapy of cancer·2026
Same author

Location patterns and longitudinal progression of white matter hyperintensities.

medRxiv : the preprint server for health sciences·2026
Same author

Incidental findings and duty-of-care protocols in cardiovascular magnetic resonance among older adults: a prospective population-based study from MyoFit46.

The lancet. Healthy longevity·2026

Related Experiment Video

Updated: May 13, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

Protein function prediction using text-based features extracted from the biomedical literature: the CAFA challenge.

Andrew Wong1, Hagit Shatkay

  • 1Computational Biology and Machine Learning Lab, School of Computing, Queen's University, Kingston, ON, K7L 3N6, Canada.

BMC Bioinformatics
|March 22, 2013
PubMed
Summary

This study introduces a text-based system (Text-KNN) for predicting protein function, outperforming baseline methods. The system shows promise for annotating uncharacterized proteins using biomedical literature.

More Related Videos

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
06:50

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions

Published on: January 26, 2024

Related Experiment Videos

Last Updated: May 13, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
16:41

A Protocol for Computer-Based Protein Structure and Function Prediction

Published on: November 3, 2011

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
06:50

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions

Published on: January 26, 2024

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Advances in sequencing have generated numerous proteins of unknown function, necessitating automated prediction systems.
  • Existing methods often rely on protein sequence or structure, but biomedical literature offers a valuable alternative feature source.
  • Previous work demonstrated the utility of text features for predicting protein subcellular location and improved performance when combined with sequence data.

Purpose of the Study:

  • To develop and evaluate a text-based system for predicting molecular function and biological process (Gene Ontology terms) for unannotated proteins.
  • To assess the performance of this system within the Critical Assessment of Function Annotations (CAFA) Challenge.
  • To compare the text-based approach against baseline methods using prior distribution and sequence similarity.

Main Methods:

  • Developed a preliminary system (Text-KNN) representing proteins with text features extracted from biomedical abstracts based on statistical properties.
  • Employed a k-nearest neighbor classifier for function prediction.
  • Utilized 5-fold cross-validation on a dataset of 36,536 proteins for training and testing.

Main Results:

  • The Text-KNN classifier achieved 62% accuracy for molecular function and 17% for biological process.
  • Text-KNN outperformed the Base-Prior classifier (43% MF, 11% BP) and showed comparable performance to Base-Seq (58% MF, 28% BP).
  • Results from the CAFA evaluation dataset are also reported.

Conclusions:

  • The text-based classifier consistently outperformed the prior distribution baseline and performed comparably to the sequence similarity baseline.
  • Combining text features with other data types may enhance prediction performance.
  • The classifier performed significantly better for predicting molecular function than biological process, a trend observed in other CAFA participants.