Related Experiment Videos
Database of repetitive elements in complete genomes and data mining using transcription factor binding sites
Jorng-Tzong Horng1, F M Lin, J H Lin
1Department of Computer Science and Information Engineering, National Central University, Jung-li City 320, Taiwan, ROC. horng@db.csie.ncu.edu.tw
Summary
Repetitive genomic elements are crucial for evolutionary genomics. This study introduces a database and data mining approach to uncover regulatory patterns within these sequences, aiding gene regulation prediction.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Repetitive elements constitute a significant portion of eukaryotic genomes, with 43% in humans and 51% in rice.
- The role of repetitive sequences in evolutionary genomics is increasingly recognized.
Purpose of the Study:
- To develop a comprehensive database of repetitive sequences.
- To apply data mining techniques to discover association rules governing transcription factor binding sites within repetitive elements.
Main Methods:
- Design and implementation of the Repeat Sequence Database (RSDB).
- Utilizing mathematical algorithms to analyze repetitive sequences (direct, inverted, palindromic).
- Applying data mining to identify association rules among transcription factor binding sites.
Main Results:
- A database containing diverse repetitive sequences was created.
- Association rules were mined, revealing patterns in binding site distribution.
- Pruning techniques identified significant associations.
Conclusions:
- Mined association rules can help identify gene classes with similar regulatory mechanisms.
- The findings facilitate accurate prediction of regulatory elements.
- The approach was validated on various genomes (C. elegans, human chromosome 22, yeast).