Related Experiment Video
Updated: Feb 16, 2026

Rapid High-throughput Species Identification of Botanical Material Using Direct Analysis in Real Time High Resolution Mass Spectrometry
Published on: October 2, 2016
THiCweed: fast, sensitive detection of sequence features by clustering big datasets.
Ankit Agrawal1, Snehal V Sambare1, Leelavati Narlikar2
1Computational Biology Group, The Institute of Mathematical Sciences (HBNI), Chennai 600113, Tamil Nadu, India.
THiCweed offers a faster and more accurate method for analyzing transcription factor binding data from ChIP-seq experiments, especially for complex datasets with mixed motifs. This new approach reveals intricate DNA sequence patterns and variant motifs, enhancing our understanding of gene regulation.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- High-throughput chromatin immunoprecipitation-sequencing (ChIP-seq) generates vast amounts of data on transcription factor binding.
- Analyzing ChIP-seq data to identify transcription factor binding motifs is crucial for understanding gene regulation.
- Traditional motif-finding tools struggle with datasets containing mixtures of motifs.
Purpose of the Study:
- To develop a novel computational approach, THiCweed, for analyzing ChIP-seq data.
- To improve the accuracy and speed of motif discovery in complex ChIP-seq datasets.
- To uncover complex sequence characteristics and recurring patterns in transcription factor binding regions.
Main Methods:
- THiCweed employs a divisive hierarchical clustering approach based on sequence similarity within sliding windows, analyzing both DNA strands.
- The method is optimized for speed, processing 30,000 peaks in 1-2 hours on a single CPU core.
- It utilizes large window sizes (≥50 bp) for analysis, differing from typical binding site lengths.
Main Results:
- THiCweed demonstrates accuracy comparable to or exceeding other programs on synthetic data with mixed motifs.
- The tool successfully identifies known motifs and discovers variant and secondary motifs present in less than 5% of the input data.
- Analysis of real ChIP-seq data revealed recurring sequence patterns potentially linked to chromatin architecture and looping.
Conclusions:
- THiCweed provides a significant advancement in analyzing ChIP-seq data, offering enhanced speed and accuracy.
- The approach moves beyond traditional motif finding to reveal complex sequence characteristics and biological insights.
- THiCweed facilitates a deeper understanding of transcription factor binding complexity and its relation to genomic architecture.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Modern Molecular Taxonomy
Evolutionary Relationships through Genome Comparisons
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Sanger Sequencing

