Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Information content of protein sequences.

O Weiss1, M A Jiménez-Montaño, H Herzel

  • 1Institute for Theoretical Biology, Humboldt University Berlin, Invalidenstr. 43, Berlin, D-10115, Germany.

Journal of Theoretical Biology
|September 16, 2000
PubMed
Summary

Protein sequences exhibit minimal redundancy, closely resembling random strings with only about 1% deviation. This finding suggests proteins are essentially edited random sequences, impacting our understanding of their complexity.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Outcomes and rate of return to play in elite athletes following arthroscopic surgery of the hip.

International orthopaedics·2021
Same author

Keeping children safe: a model for predicting families at risk for recurrent childhood injuries.

Public health·2019
Same author

Influence of wood ash pre-treatment on leaching behaviour, liming and fertilising potential.

Waste management (New York, N.Y.)·2018
Same author

Anti-phosphorylated histone H2AThr120: a universal microscopic marker for centromeric chromatin of mono- and holocentric plant species.

Cytogenetic and genome research·2014
Same author

Genetic redundancy strengthens the circadian clock leading to a narrow entrainment range.

Journal of the Royal Society, Interface·2013
Same author

Mathematical modeling in chronobiology.

Handbook of experimental pharmacology·2013

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Sequence Analysis

Background:

  • Understanding the inherent complexity and randomness of protein sequences is crucial for deciphering biological functions.
  • Previous assumptions about protein sequence complexity have not been rigorously quantified.

Purpose of the Study:

  • To quantify the complexity of large, non-redundant protein sequence datasets.
  • To determine the extent of redundancy in protein sequences compared to random sequences.

Main Methods:

  • Estimating Shannon entropy to measure sequence information content.
  • Applying compression algorithms to assess algorithmic complexity.
  • Comparing protein data complexity with randomly generated surrogate sequences.

Related Experiment Videos

Main Results:

  • Protein sequences show a low degree of redundancy, with entropy reduction around 1%.
  • Compression algorithms indicate redundancy is approximately 1%.
  • Finite sample effects limit precise entropy estimation of the source.

Conclusions:

  • Protein sequences can be accurately modeled as slightly edited random strings.
  • Observed redundancy is attributed to factors like secondary structure and low-complexity regions.
  • Findings align with experimental data from random polypeptides.