一种最大概率的方法,用SNP数据来估计单双型频率和流行率以及感染的多重性
Henri Christian Junior Tsoungui Obama1, Kristan Alexander Schneider1
1Department of Applied Computer- and Biosciences, University of Applied Sciences Mittweida, Mittweida, Germany.
Frontiers in epidemiology
|March 8, 2024
概括
这项研究引入了一个使用预期-最大化算法的统计框架,以准确估计疟疾单元型频率和流行率,即使是多次感染 (MOI). 该方法比分子疾病监测的传统方法提供了不那么有偏见的结果.
科学领域:
- 基因组流行病学基因组流行病学
- 分子疾病监测监测分子疾病的监测.
- 寄生虫学的寄生虫学
背景情况:
- 基因组方法可以实现标准化的分子疾病监测,这对于了解病原体传播动态至关重要.
- 疟疾寄生虫中的单核酸多态 (SNP) 条形码,如*Plasmodium falciparum*和*Plasmodium vivax*,有助于描述单核酸多态和传播模式.
- 感染多重性 (MOI),即在单一宿主中存在多种病原体变体,使得对遗传频率和流行率的准确估计变得复杂.
研究的目的:
- 开发一个统计框架,以获得哈普洛型频率和流行率的最大概率估计 (MLE),并考虑疟疾中的MOI.
- 为了比较使用所有遗传数据和仅使用明确数据得出的估计,突出了临时方法中的偏差.
- 为在传染病监测中分析SNP数据提供强大且数值稳定的方法.
主要方法:
- 引入使用预期-最大化 (EM) 算法来计算单型频率和流行率的MLE的统计框架.
- 由于模型参数的几何增加与遗传标记的数量,防止封闭式解决方案的MLE的数值导出.
- 使用所有遗传信息,并将其与仅从明确数据的估计值进行比较,推导出单 haplotype 的流行表达式.
主要成果:
- 使用EM算法开发的统计框架提供了数值稳定和准确的MLE,用于单种类型的频率和流行率,即使有MOI.
- 模拟表明,该方法对现实的样本大小和遗传位置的数量具有最小的偏差.
- 将其应用于来自喀麦隆的 * P. falciparum * 关于耐药性的数据集,证实了该方法的实用性.
结论:
- 拟议的统计框架有效地解决了MOI在分子监测中的挑战,产生了对哈普洛型频率和流行率的不太偏见的估计.
- 该方法比传统的临时方法显著改进,这些方法不考虑模糊的遗传信息.
- 这种方法广泛适用于任何使用类似遗传标记数据的传染病监测,并提供R脚本实现.
相关概念视频
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
Hardy-Weinberg Principle
72.1K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
72.1K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K


