Related Experiment Video
Updated: Jan 30, 2026

Quantification of Heavy Metals and Other Inorganic Contaminants on the Productivity of Microalgae
Published on: July 10, 2015
NGSTroubleFinder: a tool for detection and quantification of contamination and kinship across human NGS data
Samuel Valentini1, Tecla Venturelli1, Xavier Gallego1
1STALICLA Discovery and Data Science Unit, World Trade Center, Moll de Barcelona, Edif Este, 08039 Barcelona, Spain.
Abstract:
Quality control constitutes a critical component of any next-generation sequencing (NGS) pipeline; however, most existing pipelines emphasize technical quality assessment (e.g. read quality, alignment metrics, duplication rates) while overlooking other equally important dimensions, such as sample identity verification, contamination detection, kinship analysis, and metadata concordance. Detecting issues like cross-sample contamination and sample swaps is essential to control data integrity. Here, we present NGSTroubleFinder, a novel tool to detect cross-sample contamination in human whole-genome and whole-transcriptome sequencing data, sample swaps, and mismatches between the reported and the inferred genetic and transcriptomic sexes. It can be run directly on BAM/CRAM files without requiring additional variant-calling steps and offers an integrated pipeline for ensuring quality control on NGS data, generated particularly within the context of clinical studies or research projects involving family members. It produces a detailed report that combines the results of its multiple analyses, including kinship, sex prediction, and contamination metrics. The tool reports extensive information on the samples, both in textual and HTML formats, including key plots for easy interpretation of the results. NGSTroubleFinder is written in Python and incorporates a custom-built parallelized pileup engine written in C, and it can be easily installed with pip. The tool source code and the models are freely available on GitHub (https://github.com/STALICLA-RnD/NGSTroubleFinder), and a containerized version is available on Docker Hub (https://hub.docker.com/r/staliclarnd/ngstroublefinder).
More Related Videos
Related Concept Videos
Overview of Microsoft Excel as a Data Analysis Tool
Contaminants and Errors
Another key consideration is determining the appropriate number of samples required to...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Data Reporting and Recording
Data Validation
Key parameters for method validation include:

