Related Experiment Videos
Fast identification of repetitive elements in biological sequences
1Laboratoire de Chimie Bactérienne, Centre National de la Recherche Scientifique, Marseille, France.
Journal of Theoretical Biology
|January 7, 1994
Summary
A new k-word frequency analysis method efficiently identifies repetitive DNA sequences, like Alu elements, in large genomic datasets. This tool aids in discovering unannotated repetitive elements for systematic database screening and specialized database creation.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Repetitive DNA sequences are abundant in genomes but often lack annotation in databases.
- Efficient methods are needed to identify and classify these repetitive elements for further analysis.
Purpose of the Study:
- To develop a fast filtering method for simultaneously identifying diverse families of repetitive elements in biological sequences.
- To enable systematic screening of new sequences for repetitive elements before database submission.
Main Methods:
- A k-word frequency comparison method was developed to discriminate between repetitive and non-repetitive sequences.
- Correspondence analysis was used to weight k-words for improved sequence discrimination.
- The method was tested for identifying Alu elements in human DNA sequences.
Main Results:
- The method achieved high accuracy, with only 0.5% misclassification of non-Alu sequences and 1.4% misclassification among Alu monomers.
- All Alu elements (616) were correctly identified in 63 large GenBank sequences, with minimal misprediction of non-Alu fragments (22).
- A word length of 6 base pairs provided excellent discrimination.
Conclusions:
- The developed method offers a fast and accurate approach for identifying repetitive DNA elements.
- This tool is valuable for annotating uncharacterized repetitive sequences and creating specialized databases.
- It facilitates systematic genomic analysis and improves the quality of public sequence databases.