我们会丢弃好的数据吗? 在长时间读取的安普利康上对奇梅拉检测算法的评估显示,在算法中,错误阳性率很高
Ali Hakimzadeh1, Vladimir Mikryukov1,2, Martin Metsoja1
1Institute of Ecology and Earth Sciences, University of Tartu, Tartu, Estonia.
PeerJ
|December 10, 2025
概括
长时间读取的安普利康序列的基准化默拉检测算法显示,uchime_denovo提供了卓越的精度. 虽然大多数算法都在与假阳性作斗争,但整体生物多样性估计在各种方法中保持稳健.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 生态生态学 生态生态学
背景情况:
- 长时间读取的安普利康序列增强了使用DNA条形码的生物多样性研究中的分类学分辨率.
- 在PCR过程中形成的化学序列,通过潜在地扭曲多样性估计和生态解释,构成重大挑战.
研究的目的:
- 为了对三种新兴的奇默探测算法进行基准测试 (uchime_denovo, removeBimeraDenovo, chimeras_denovo).
- 通过模拟和实证的真核生物全ITS数据集,评估它们的精度,敏感性和对操作分类学单位 (OTU) 和社区结构的影响.
主要方法:
- 用 uchime_denovo, removeBimeraDenovo 和 chimeras_denovo 算法进行基准测试.
- 使用模拟和实证真核生物全ITS (rRNA ITS1-5.8S-ITS2) 序列数据.
- 评估了准确性,敏感性和对OTU组成和社区结构的影响.
主要成果:
- uchime_denovo在模拟数据上展示了最高的精度,即使具有默认设置,与其他算法不同,需要进行调整以最大限度地减少假阳性.
- 经验数据分析证实了 uchime_denovo 的较低假阳性率; chimeras_denovo 和 removeBimeraDenovo 错误地将大约一半的序列标记为虚构的.
- 假阴性仿真体通常包含多个5.8S区域,表明图书馆准备文物 (PacBio) 而不是PCR文物.
结论:
- uchime_denovo 是一个更精确的工具,用于在长读 amplicon 测序中检测金像.
- 尽管奇美拉检测挑战和潜在的文物,总体生物多样性指标和社区模式在不同过策略中基本一致.
相关概念视频
Mismatch Repair
Overview
Mismatch Repair
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Systematic Error: Methodological and Sampling Errors
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Quantifying and Rejecting Outliers: The Grubbs Test
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Data Validation
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
Key parameters for method validation include:


