Related Experiment Video
Updated: Nov 17, 2025

11:41
Mapping Mammalian 3D Genome Interactions with Micro-C-XL
Published on: November 3, 2023
3.1K
S-conLSH: alignment-free gapped mapping of noisy long reads
Angana Chakraborty1, Burkhard Morgenstern2, Sanghamitra Bandyopadhyay3
1Department of Computer Science, West Bengal Education Service, Kolkata, India.
BMC Bioinformatics
|February 12, 2021
Summary
A new alignment-free sequence mapping tool, S-conLSH, offers fast and accurate genome analysis for long reads. This method excels at identifying distant homologies, improving genome analysis pipelines.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Single Molecule, Real-Time (SMRT) sequencing offers long reads and low GC bias, advancing genome analysis.
- Traditional alignment methods struggle with distant homologies in genomes due to genetic duplication and recombination.
- Accurate read mapping is a critical and computationally intensive step in SMRT sequencing analysis.
Purpose of the Study:
- To develop a novel, fast, and accurate alignment-free method for mapping long sequencing reads.
- To overcome the limitations of existing aligners in identifying distant homologies in complex genomes.
Main Methods:
- Introduced S-conLSH, a novel mapper utilizing Spaced context based Locality Sensitive Hashing.
- Employed multiple spaced patterns for gapped mapping of noisy long reads to reference genomes.
- Evaluated performance on five real and simulated datasets, comparing against existing methods like lordFAST.
Main Results:
- S-conLSH demonstrated at least a 2-fold increase in speed compared to lordFAST.
- Achieved 99% sensitivity on human simulated sequence data without traditional base-to-base alignment.
- Provides alignment-free mapping in PAF format by default, with an option for SAM-file output.
Conclusions:
- S-conLSH is a pioneering alignment-free tool for reference genome mapping with high sensitivity.
- The spaced-context approach effectively extracts distant similarities, crucial for complex genomic regions.
- Variable-length spaced patterns enhance flexibility for gapped mapping of noisy long reads, advancing alignment-free analysis.
Related Concept Videos
Sanger Sequencing
766.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
766.3K
RNA-seq
11.1K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.1K
Genome Copying Errors
4.8K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.8K
Improving Translational Accuracy
12.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
12.5K
Improving Translational Accuracy
3.3K
3.3K
Mismatch Repair
5.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
5.8K

