Related Experiment Videos
Information content of protein sequences
O Weiss1, M A Jiménez-Montaño, H Herzel
1Institute for Theoretical Biology, Humboldt University Berlin, Invalidenstr. 43, Berlin, D-10115, Germany.
Journal of Theoretical Biology
|September 16, 2000
Summary
Protein sequences exhibit minimal redundancy, closely resembling random strings with only about 1% deviation. This finding suggests proteins are essentially edited random sequences, impacting our understanding of their complexity.
Area of Science:
- Bioinformatics
- Computational Biology
- Sequence Analysis
Background:
- Understanding the inherent complexity and randomness of protein sequences is crucial for deciphering biological functions.
- Previous assumptions about protein sequence complexity have not been rigorously quantified.
Purpose of the Study:
- To quantify the complexity of large, non-redundant protein sequence datasets.
- To determine the extent of redundancy in protein sequences compared to random sequences.
Main Methods:
- Estimating Shannon entropy to measure sequence information content.
- Applying compression algorithms to assess algorithmic complexity.
- Comparing protein data complexity with randomly generated surrogate sequences.
Main Results:
- Protein sequences show a low degree of redundancy, with entropy reduction around 1%.
- Compression algorithms indicate redundancy is approximately 1%.
- Finite sample effects limit precise entropy estimation of the source.
Conclusions:
- Protein sequences can be accurately modeled as slightly edited random strings.
- Observed redundancy is attributed to factors like secondary structure and low-complexity regions.
- Findings align with experimental data from random polypeptides.