相关实验视频
Updated: Feb 7, 2026

08:57
Amplicon Sequencing using the Long-Read Sequencing Technologies
Published on: August 29, 2025
537
异常:Snakemake管道用于从长时间读取的测序数据中识别NuMT
Nirmal S Mahar1, Rachit Singh2, Ishaan Gupta1
1Department of Biochemical Engineering and Biotechnology, Indian Institute of Technology Delhi, New Delhi 110016, India.
NAR genomics and bioinformatics
|February 6, 2026
概括
核线粒体DNA片段 (NuMTs) 可以驱动癌症. 一个新的工作流,ANOMALY,使用长读序列精确检测NuMTs,克服了旧方法的局限性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 癌症研究 癌症研究
背景情况:
- 核线粒体DNA片段 (NuMTs) 与癌症发展有关,并使线粒体变异分析复杂化.
- 使用短读序列数据的现有检测方法不足以解决复杂的NuMTs.
研究的目的:
- 开发和验证一个新的工作流程,以从长时间读取的测序数据中准确检测NuMT.
- 为了解决当前NuMT检测方法的局限性.
主要方法:
- 介绍ANOMALY,一个易于使用的工作流程,用于从长读序列数据中检测NuMT.
- 管道处理原始或对齐的测序数据以识别和可视化NuMTs.
- 使用Python,Bash和R,基于Snakemake框架构建.
主要成果:
- 在50个模拟数据集上,ANOMALY表现出高精度,精度为1.000,回忆率为0.989和F1得分为0.994.
- 与短读数据相比,证实了长读测序数据在解析和捕获复杂的NuMTs方面的优势.
- 该工作流是在开源GNU GPL v3许可证下发布的,代码可在GitHub上找到.
结论:
- 通过ANOMALY处理的长读序列数据,可以准确识别NuMTs.
- 异常工作流提供了一个强大的解决方案,用于NuMT检测,这对于癌症研究和诊断至关重要.
- 突出了复杂的基因组结构变异的短读测序的局限性.
相关概念视频
Uncertainty in Measurement: Reading Instruments
53.1K
Counting is the type of measurement that is free from uncertainty, provided the number of objects being counted does not change during the process. Such measurements result in exact numbers. By counting the eggs in a carton, for instance, one can determine exactly how many eggs are there in the carton. Similarly, the numbers of defined quantities are also exact. For example, 1 foot is exactly 12 inches, 1 inch is exactly 2.54 centimeters, and 1 gram is exactly 0.001 kilograms. Quantities...
53.1K
Cis-regulatory Sequences
11.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.9K
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K
How Data are Classified: Numerical Data
38.1K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.1K
Sequences
280
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
280
Sanger Sequencing
774.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.7K

