一个改进的框架,用于检测分离的流行病学上有意义的分区在等级分类的遗传数据中
David K Jacobson1, Ross Low1,2, Mateusz M Plucinski1
1Division of Parasitic Diseases and Malaria, Centers for Disease Control and Prevention, Atlanta, GA, 30329, United States.
Bioinformatics advances
|September 25, 2023
概括
方法B准确地剖析了层级的微生物基因型集群,反映了食物传播寄生虫的流行病学分组. 与其他树切割方法相比,这种框架提供了更高的准确性和自动化值.
科学领域:
- 微生物学 微生物学
- 计算生物学 计算生物学
- 流行病学 流行病学
背景情况:
- 微生物基因型的等级聚类会产生嵌套的聚类,这对流行病学解释提出了挑战.
- 将这些树切割成有意义的分组需要方法,以尽量减少调查者偏见.
- 之前的工作引入了基于统计框架的树剖析方法A.
研究的目的:
- 应用一个修改后的框架,方法B,用于剖析Cyclospora基因型的等级树.
- 评估B方法与流行病学定义的疫情爆发集群的性能.
- 为了将方法B与现有的树切割方法进行比较:方法A,cutreeHybrid,cutreeDynamic,TreeCluster和PARNAS.
主要方法:
- 应用方法B对211个Cyclospora基因型的数据集,包括639个与疫情相关的病例.
- 评估了B方法的遗传分区与已知的流行病学集群之间的一致性.
- 与使用流行病学数据的其他树剖析算法比较方法B的准确性和歧视性.
主要成果:
- 方法B在识别反映流行病学分类的遗传群体方面取得了99.4%的准确性,类似于TreeCluster和PARNAS.
- CutreeHybrid显示出更广泛的种群结构,但缺乏菌株水平的细节;cutreeDynamic具有良好的菌株歧视,但敏感性较低.
- 方法B自动计算树切割值,以默认值优于TreeCluster.
结论:
- 方法B提供了一种准确和自动化的方法来剖析层次的微生物基因型集群.
- 该框架有效地确定了流行病学相关的分组,这对于了解病原体传播至关重要.
- 方法B为微生物流行病学提供了有价值的工具,具有公开可用的代码和数据.
相关概念视频
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Statistical Methods for Analyzing Epidemiological Data
400
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
400
Sampling Plans
205
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
205
Statistical Software for Data Analysis and Clinical Trials
610
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
610


