Related Experiment Videos
Automated biological sequence description by genetic multiobjective generalized clustering
I Zwir1, R Romero Zaliz, E H Ruspini
1Department of Molecular Microbiology, Howard Hughes Medical Institute Research Laboratories, Washington University School of Medicine, Saint Louis, Missouri 63110-1093, USA. zwir@borcim.wustl.edu
Annals of the New York Academy of Sciences
|February 21, 2003
Summary
This study introduces a novel method for identifying significant qualitative features in biological sequences, using a generalized clustering approach. This helps in understanding complex biological data and discovering patterns in DNA sequences.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Databases for complex biological objects are advancing, but tools for retrieving and understanding them are lacking.
- Current methods for DNA sequence analysis struggle to identify features based on qualitative characteristics.
Purpose of the Study:
- To develop a method for identifying interesting qualitative features in biological sequences.
- To improve the retrieval and understanding of complex biological data.
Main Methods:
- A generalized clustering methodology is used, framing feature identification as a multivariable, multiobjective optimization problem.
- Genetic algorithms solve the optimization problem, discovering candidate features as fuzzy subsets.
- Evolutionary computation methods summarize and relate these features.
Main Results:
- The method successfully identifies and summarizes interesting features in DNA sequences.
- Applied to Tripanosoma cruzi DNA, it demonstrated recognition and summarization of significant features.
Conclusions:
- The proposed two-step method effectively addresses the need for qualitative feature identification in biological sequences.
- This approach enhances the analysis and interpretation of complex biological datasets, particularly DNA sequences.