Related Experiment Video
Updated: Jun 5, 2026

Immunostaining for DNA Modifications: Computational Analysis of Confocal Images
Published on: September 7, 2017
WordCluster: detecting clusters of DNA words and genomic elements
Michael Hackenberg1, Pedro Carpena, Pedro Bernaola-Galván
1Dpto, de Genética, Facultad de Ciencias, Universidad de Granada, Campus de Fuentenueva s/n, 18071-Granada & Lab, de Bioinformática, Centro de Investigación Biomédica, PTS, Avda, del Conocimiento s/n, 18100-Granada, Spain. mlhack@gmail.com.
A new algorithm, WordCluster, identifies statistically significant clusters of DNA words (k-mers) and genomic elements. This tool aids in understanding genome organization and function, revealing biological insights like methylation patterns.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Genomic elements and k-mers (DNA words) exhibit spatial clustering, evident in genes, transcription factor binding sites (TFBSs), and CpG dinucleotides.
- Existing methods for detecting genomic clusters often lack statistical rigor, relying on density, sliding-window approaches, or arbitrary distance thresholds.
Purpose of the Study:
- To introduce a novel algorithm, WordCluster, for statistically sound detection of clusters of k-mers and other genomic elements.
- To develop a web server implementation of the algorithm for user accessibility and integration with genomic annotation data.
Main Methods:
- Developed an algorithm that detects clusters based on the distance between consecutive k-mer copies and assigns statistical significance.
- Implemented the algorithm into a web server with a MySQL backend for co-localization analysis with gene annotations.
- Utilized the tool to identify clusters of CAG/CTG repeats and olfactory receptor (OR) genes in the human genome.
Main Results:
- Successfully detected statistically significant clusters of CAG/CTG repeats, revealing significant variations in methylation levels inside and outside these clusters.
- Identified statistically significant clusters of olfactory receptor (OR) genes in the human genome, demonstrating the tool's applicability to gene families.
- The WordCluster algorithm proves effective in identifying biologically meaningful clusters of genomic entities.
Conclusions:
- WordCluster accurately predicts biologically relevant clusters of DNA words and genomic entities.
- The web server implementation offers enhanced features, including co-localization with gene regions and functional annotation enrichment analysis.
- The tool is publicly available at http://bioinfo2.ugr.es/wordCluster/wordCluster.php for broader research use.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
DNA Microarrays
Labeling DNA Probes
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Genomic DNA in Eukaryotes
