Related Experiment Video
Updated: Feb 7, 2026

08:57
Amplicon Sequencing using the Long-Read Sequencing Technologies
Published on: August 29, 2025
537
ANOMALY: a Snakemake pipeline for identifying NuMTs from long-read sequencing data.
Nirmal S Mahar1, Rachit Singh2, Ishaan Gupta1
1Department of Biochemical Engineering and Biotechnology, Indian Institute of Technology Delhi, New Delhi 110016, India.
NAR Genomics and Bioinformatics
|February 6, 2026
Summary
Nuclear mitochondrial DNA segments (NuMTs) can drive cancer. A new workflow, ANOMALY, accurately detects NuMTs using long-read sequencing, overcoming limitations of older methods.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- Nuclear mitochondrial DNA segments (NuMTs) are implicated in cancer development and complicate mitochondrial variant analysis.
- Existing detection methods using short-read sequencing data are insufficient for resolving complex NuMTs.
Purpose of the Study:
- To develop and validate a novel workflow for accurate NuMT detection from long-read sequencing data.
- To address the limitations of current NuMT detection methods.
Main Methods:
- Introduction of ANOMALY, an easy-to-use workflow for NuMT detection from long-read sequencing data.
- The pipeline processes raw or aligned sequencing data to identify and visualize NuMTs.
- Utilizes Python, Bash, and R, built on the Snakemake framework.
Main Results:
- ANOMALY demonstrated high accuracy on 50 simulated datasets, achieving precision of 1.000, recall of 0.989, and an F1-score of 0.994.
- Confirmed the superiority of long-read sequencing data for resolving and capturing complex NuMTs compared to short-read data.
- The workflow is published under an open-source GNU GPL v3 license with code available on GitHub.
Conclusions:
- Long-read sequencing data, processed by ANOMALY, enables accurate identification of NuMTs.
- The ANOMALY workflow provides a robust solution for NuMT detection, crucial for cancer research and diagnostics.
- Highlights the limitations of short-read sequencing for complex genomic structural variations.
Related Concept Videos
Uncertainty in Measurement: Reading Instruments
53.1K
Counting is the type of measurement that is free from uncertainty, provided the number of objects being counted does not change during the process. Such measurements result in exact numbers. By counting the eggs in a carton, for instance, one can determine exactly how many eggs are there in the carton. Similarly, the numbers of defined quantities are also exact. For example, 1 foot is exactly 12 inches, 1 inch is exactly 2.54 centimeters, and 1 gram is exactly 0.001 kilograms. Quantities...
53.1K
Cis-regulatory Sequences
11.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.9K
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K
How Data are Classified: Numerical Data
38.1K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.1K
Sequences
280
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
280
Sanger Sequencing
774.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.7K

