Related Experiment Videos
Discovering patterns and subfamilies in biosequences
A Brazma1, I Jonassen, E Ukkonen
1Institute of Mathematics and Computer Science, University of Latvia, Riga, Latvia. abrozma@cclu.lv
Summary
This study introduces a novel method for simultaneously discovering patterns and subfamilies in unaligned biosequences. The approach, based on the minimum description length principle, accurately identifies protein family patterns and structures.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biological sequence analysis often requires identifying conserved patterns and grouping related sequences into subfamilies.
- Existing methods may struggle with unaligned sequences or noise, limiting their effectiveness.
- The PROSITE database provides a valuable resource for protein family patterns.
Purpose of the Study:
- To develop a method for the simultaneous automatic discovery of patterns and subfamilies in unaligned biosequences.
- To address challenges posed by noisy and unaligned sequence data.
- To create a theoretically grounded significance measure for discovered patterns.
Main Methods:
- Utilizing the minimum description length (MDL) principle for pattern and subfamily discovery.
- Developing an algorithm to approximate the optimal set of patterns and corresponding subfamilies.
- Applying the method to unaligned biosequences with unknown noise levels.
Main Results:
- Successfully identified subfamilies within the chromo domain protein family.
- Discovered novel and significant patterns within the analyzed biosequence dataset.
- Demonstrated the effectiveness of the MDL-based approach in handling noisy sequence data.
Conclusions:
- The developed method provides a robust framework for simultaneous pattern and subfamily discovery in bioinformatics.
- The MDL principle offers a theoretically sound basis for evaluating pattern significance.
- This approach enhances the ability to analyze and classify protein families, aiding in functional annotation.