Related Experiment Video
Updated: Jun 7, 2026

16:17
The ITS2 Database
Published on: March 12, 2012
30.8K
ntsm: an alignment-free, ultra-low-coverage, sequencing technology agnostic, intraspecies sample comparison tool for
Justin Chu1,2, Jiazhen Rong3, Xiaowen Feng1,2
1Dana-Farber Cancer Institute, Department of Data Sciences, Boston, MA 02215, USA.
Gigascience
|June 4, 2024
Summary
A new method detects sample swaps in large genomic studies using k-mer variants and coverage data, improving quality control and saving computational resources across diverse sequencing data types.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Sample swapping is a common error in large cohort studies, particularly with heterogeneous data types from various sequencing technologies.
- Current sample swap detection methods are computationally expensive, requiring data alignment, sorting, and indexing, which is often unnecessary for downstream analyses like genome assembly.
- The increasing scale and complexity of genomic studies necessitate robust and efficient quality control tools.
Purpose of the Study:
- To develop a faster and more resource-efficient method for detecting sample swaps in large-scale genomic studies.
- To provide a quality control tool applicable to diverse sequencing data types without requiring data alignment.
- To offer additional QC metrics such as error rates and coverage bias.
Main Methods:
- Utilizes indexed k-mer sequence variants to determine sample similarity.
- Incorporates coverage information at variant sites to enhance statistical power via a likelihood ratio-based test.
- Employs spatially indexed principal component analysis (PCA)-based prescreening to accelerate analysis by avoiding exhaustive comparisons.
Main Results:
- The method accurately identifies sample swaps by analyzing k-mer variants and coverage data.
- It enables estimation of per-sample error rates and coverage bias.
- PCA-based prescreening significantly speeds up the analysis process.
Conclusions:
- This novel tool processes raw sequencing data, bypassing the need for alignment, thus saving significant computational resources in standard quality control (QC) pipelines.
- The method is robust across different sequencing data types (e.g., Oxford Nanopore, Pacific Biosciences, Illumina), making it suitable for studies combining multiple technologies.
- Beyond sample swap detection, it provides valuable QC information and facilitates population-level PCA for ancestry analysis visualization.

