Related Experiment Video
Updated: Dec 25, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
Binning unassembled short reads based on k-mer abundance covariance using sparse coding
Olexiy Kyrgyzov1, Vincent Prost1,2, Stéphane Gazut2
1Génomique Métabolique, Genoscope, Institut François Jacob, CEA, CNRS, Université Paris-Saclay, 2 rue Gaston Crémieux, 91057 Evry, France.
This study introduces a scalable pre-assembly binning method for recovering microbial genomes from metagenomes. The novel approach, using sparse dictionary learning, successfully identifies low-abundance genomes without prior assembly, complementing existing methods.
Area of Science:
- Microbiology
- Bioinformatics
- Computational Biology
Background:
- Metagenome assembly is computationally intensive and can miss low-abundance genomes.
- Existing sequence-binning techniques often rely on prior metagenome assembly.
- Large-scale metagenomic datasets present computational challenges for assembly.
Purpose of the Study:
- To develop a scalable pre-assembly binning scheme for microbial genome recovery.
- To enable the recovery of low-abundance genomes from complex metagenomes.
- To leverage sparse dictionary learning for read-level binning.
Main Methods:
- A pre-assembly binning scheme operating on unassembled short reads.
- Application of sparse dictionary learning and elastic-net regularization.
- Joint analysis of microbiomes from the LifeLines DEEP population cohort (n = 1,135).
Main Results:
- Recovery of hundreds of metagenome-assembled genomes, including very low-abundance ones.
- Demonstration of read-level binning at scale using sparse coding techniques.
- Observed read enrichment across six orders of magnitude in relative abundance.
Conclusions:
- Sparse coding enables scalable read-level binning for microbial genome recovery.
- Bin-first strategies can complement assembly-first protocols by targeting distinct genome segregation.
- The method effectively recovers genomes with low relative abundance.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Sanger Sequencing

