Related Experiment Videos
A new family of powerful multivariate statistical sequence analysis techniques.
1Fritz Haber Institute of the Max Planck Society, Berlin Dahlem, Germany.
Journal of Molecular Biology
|August 20, 1991
Summary
A new statistical method, Multivariate Statistical Sequence Analysis (MSSA), analyzes sequence data without alignment. This approach reveals biological details, aiding in tasks like phylogenetic comparisons and genome projects.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- Sequence data banks are growing rapidly.
- Traditional sequence alignment methods have limitations.
Purpose of the Study:
- To present a novel multivariate statistical approach for sequence data analysis.
- To extract intrinsic information from sequences without intersequence alignment.
Main Methods:
- Analyzing secondary invariant functions derived from sequences.
- Using a 20x20 histogram of amino acid pair occurrences as a typical invariant function.
- Applying Multivariate Statistical Sequence Analysis (MSSA) principles.
Main Results:
- Analysis of 10,000 protein sequences revealed significant biological detail.
- Identified evolutionary relationships, e.g., zeta-hemoglobin proximity to amphibian and fish chi-hemoglobin.
- Demonstrated unification of phylogenetic comparisons and distance matrices.
Conclusions:
- MSSA offers a robust framework for diverse sequence analysis tasks.
- Applications include family membership assignment, sequence validation, structure prediction, and coding region discrimination.
- MSSA is well-suited for continuous learning from growing sequence data and major projects like the Human Genome Project.