Related Experiment Videos
Distribution patterns of over-represented k-mers in non-coding yeast DNA
Steven Hampson1, Dennis Kibler, Pierre Baldi
1Department of Information and Computer Science, Institute for Genomics and Bioinformatics, University of California, Irvine, Irvine, CA 92697-3425, USA. hampson@ics.uci.edu
Bioinformatics (Oxford, England)
|May 23, 2002
Summary
We identified non-random spatial distributions of over-represented k-mers in yeast upstream regions. These patterns reveal insights into DNA structure, function, and evolution, aiding in the discovery of regulatory elements.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Over-represented k-mers in genomic DNA can indicate biologically significant regions.
- These k-mers are often associated with transcription factor binding sites in co-regulated gene families.
- Understanding k-mer distribution is crucial for deciphering regulatory elements.
Purpose of the Study:
- To introduce a statistical background model for measuring k-mer over-representation.
- To investigate the context and spatial distribution of over-represented k-mers in yeast upstream regions.
- To relate k-mer distribution patterns to DNA structure, function, and evolution.
Main Methods:
- Development of a statistical background model based on single-mismatches.
- Application of the model to pooled 500 bp ORF Upstream Regions (USRs) of yeast.
- Analysis of spatial distributions, homology, localization, and DNA structure of over-represented k-mers.
Main Results:
- Most over-represented k-mers exhibit non-random spatial distributions, clustering into distinct classes.
- Three common patterns were identified: localized distributions in homologous ORF clusters, broad symmetric distributions for regulatory elements, and broad distributions for A/T-rich runs.
- These patterns correlate with DNA structure, function, and evolutionary relationships.
Conclusions:
- The spatial distribution of over-represented k-mers provides valuable information about their biological roles.
- A data-mining approach integrating over-representation, homology, localization, and DNA structure aids in identifying important k-mers.
- This study contributes to understanding the 'lexicon' of regulatory regions in genomic DNA.